DEV Community

Cover image for OnTrack: Real-Time Agent Monitoring via Streaming Optimal Transport
mech.app
mech.app

Posted on Originally published at mech.app

OnTrack: Real-Time Agent Monitoring via Streaming Optimal Transport

Autonomous agents execute irreversible actions. A stock trading agent places orders. A trip planner books flights. An IT triage agent restarts production services. Current safeguards are either too slow (post-hoc log analysis) or too expensive (a second LLM watching every step). OnTrack introduces a third option: streaming trajectory comparison that flags divergence in about a millisecond per step, before the damage compounds.

The core idea is structure-aware optimal transport. Instead of waiting for task completion or running a shadow agent, OnTrack compares the unfolding trajectory against reference runs in real time. It tracks not just what tools are called, but the dependency graph between steps. When an agent starts looping, stalling, or deviating from known-good patterns, the system raises an alert or blocks execution.

The Intervention Window Problem

Post-hoc observability tells you what went wrong after tokens are burned and APIs are hit. Safeguard agents (a second LLM evaluating each step) add 200-500ms of latency and double your inference cost. Neither works for high-frequency trading agents or cost-sensitive production deployments.

OnTrack targets the gap between action and consequence:

  • Latency budget: ~1ms per step for trajectory comparison
  • Intervention trigger: Structural divergence from reference trajectories
  • Failure modes caught: Loops, stalls, tool call repetition, plan violations

The system operates in three regimes based on available data:

Access Level Available Data Detection Capability
Full Historical runs + tool schemas Plan violation, dependency graph mismatch
Intermediate Tool schemas only Loop detection, repeated tool calls
Minimal Step logs as generated Stall detection, basic anomaly flagging

Streaming Optimal Transport Architecture

Traditional optimal transport computes the minimum cost to transform one distribution into another. OnTrack adapts this for streaming agent trajectories by comparing partial execution graphs in real time.

State representation:

  • Each step is a node: tool call, reasoning step, or API response
  • Edges capture dependencies: which outputs feed which inputs
  • Graph structure encodes the agent's plan, not just the sequence

Comparison mechanism:

  1. Reference trajectories are pre-processed into canonical dependency graphs
  2. As the agent executes, OnTrack builds the current trajectory graph incrementally
  3. Optimal transport distance measures structural divergence at each step
  4. Threshold crossing triggers intervention (alert or block)

The structure-aware component is critical. Two agents might call the same tools in different orders but produce equivalent results. OnTrack's graph comparison distinguishes between benign reordering and dangerous divergence.

# Simplified trajectory comparison pseudocode
class TrajectoryMonitor:
    def __init__(self, reference_graphs, threshold=0.3):
        self.references = reference_graphs
        self.threshold = threshold
        self.current_graph = DependencyGraph()

    def observe_step(self, step):
        # Add step to current trajectory graph
        self.current_graph.add_node(step)

        # Compute streaming OT distance to nearest reference
        min_distance = min(
            optimal_transport_distance(self.current_graph, ref)
            for ref in self.references
        )

        # Trigger intervention if divergence exceeds threshold
        if min_distance > self.threshold:
            return InterventionSignal(
                distance=min_distance,
                action="BLOCK" if min_distance > 0.5 else "ALERT"
            )

        return None
Enter fullscreen mode Exit fullscreen mode

Reference Trajectory Construction

The system needs examples of successful runs. For stock trading agents, this creates a bootstrapping problem: what counts as a "safe" reference when optimal strategies are contested?

Three approaches:

  1. Supervised curation: Domain experts label successful trajectories and annotate acceptable variations
  2. Outcome-based filtering: Use trajectories that met success criteria (profitable trades, completed bookings, resolved incidents)
  3. Constraint satisfaction: Define hard boundaries (never exceed position limits, always check balance before trade) and accept any trajectory that respects them

The paper evaluates OnTrack on SWE-bench, where success is well-defined (tests pass). In financial domains, you need explicit policy: is a 2% loss acceptable exploration or a failure to intervene?

Intervention Policy and False Positives

Blocking an agent mid-execution has consequences. If the agent was exploring a novel but valid solution path, intervention wastes the partial work. If it was heading toward a catastrophic action, blocking saves money and reputation.

OnTrack's evaluation on SWE-bench shows the trade-off:

  • After 8 steps, the system ranks failing trajectories below successful ones with +0.057 AUROC improvement over content similarity
  • An abort policy saves ~18% of compute on failing runs
  • 83% precision: 5 out of 6 aborted runs were actually heading to failure

For financial agents, you tune the threshold based on risk tolerance:

  • Conservative (threshold=0.2): Catch more failures, higher false positive rate, more human review
  • Permissive (threshold=0.5): Allow exploration, miss some failures, lower operational overhead

Deployment Shape

OnTrack runs as a sidecar process or inline middleware in the agent orchestration layer. It does not replace your LLM or tool execution runtime.

Integration points:

  • Pre-action hook: Before each tool call, pass the proposed step to OnTrack
  • Streaming graph update: OnTrack maintains the dependency graph in memory
  • Intervention callback: On threshold breach, invoke your policy (log, alert, block, request human approval)

Latency characteristics:

  • Graph update: <0.5ms per step
  • OT distance computation: <1ms for graphs with <100 nodes
  • Total overhead: ~1-2ms per agent step

For high-frequency trading, this fits within the execution budget. For trip planning or IT triage, it's negligible.

Observability and Debugging

When OnTrack blocks an agent, you need to understand why. The system exposes:

  • Divergence heatmap: Which steps contributed most to the OT distance
  • Reference trajectory diff: Visual comparison of current vs. nearest reference graph
  • Counterfactual projection: If the agent had continued, which reference would it have converged toward (if any)

This turns intervention from a black-box block into a debugging signal. If OnTrack repeatedly flags a new tool usage pattern that succeeds, you update your reference set.

Failure Modes

Stale references: If your agent's environment changes (new APIs, updated tool schemas), old reference trajectories become misleading. Solution: continuous reference set refresh from recent successful runs.

Adversarial agents: A sufficiently capable agent could learn to mimic reference trajectory structure while pursuing a different goal. OnTrack is not a security boundary; it's a safety rail.

High-variance domains: In creative tasks (content generation, exploratory research), there may be no canonical "good" trajectory. OnTrack works best when success has structural invariants.

Cold start: With zero reference trajectories, the system can only detect loops and stalls. You need at least a handful of successful runs to enable plan violation detection.

Technical Verdict

Use OnTrack when:

  • Your agents execute irreversible actions (trades, bookings, infrastructure changes)
  • You have examples of successful runs to build reference trajectories
  • Post-hoc analysis is too late and safeguard agents are too expensive
  • Structural divergence (loops, plan violations) is a meaningful failure signal

Avoid OnTrack when:

  • Your domain has high variance in valid solution paths (creative tasks, open-ended research)
  • You lack reference trajectories and cannot define structural success criteria
  • Your agents operate in rapidly changing environments where references go stale quickly
  • Latency budget is so tight that even 1ms per step is unacceptable

For financial trading agents, trip planners, and IT triage, OnTrack offers a practical middle ground: real-time intervention without the cost of a shadow LLM. The key is curating reference trajectories that capture your risk boundaries, not just historical success.


Source Links

Top comments (0)