DEV Community

Ashwin Ugale
Ashwin Ugale

Posted on Originally published at arize.com

What agent traces can tell you without an LLM judge

TL;DR

  • Some agent failures can be proven from the trace alone, such as a tool call that violates its schema or a failed result reused in a side effect. Other patterns, like repeated calls with no progress, can be surfaced as candidates without claiming they're definitely bugs.
  • A useful rule for deciding which is which: if the verdict depends on what the agent should have meant, it's an eval. If it depends only on what the trace records, it can be a deterministic check.
  • tracelint is a small open-source linter that runs those checks on OpenInference traces, including the ones Phoenix collects, and returns an exit code you can use in CI.

A run can look successful and still have a broken trajectory

In a small experiment, a LangGraph release agent (gpt-4o-mini) is asked to ship release 4.12.0. Its stubbed pipeline tool returns HTTP 200 with a Jenkins result of UNSTABLE (3 of 215 tests failed). The agent deploys that build to production anyway and replies:

The release version 4.12.0 has been shipped to production. However, please note that the Jenkins build result was UNSTABLE, with 3 test failures out of 215 tests.

Every span reports success. The agent noticed the failure, and shipped anyway.

With "deploy only if its pipeline passes" in the system prompt, it never deployed in ten runs. With that sentence removed, it deployed in all ten. That's the kind of regression CI exists to catch.

Do you need an LLM judge here? No. The pipeline failed by your team's own definition, the same build ID shows up in the deploy call, and deploying has side effects. The verdict is already in the trace, plus two facts about your tools. I built tracelint to turn failures like this into repeatable checks you can run locally or in CI.

What can be proven from a trace?

The rule I use is this. If the verdict requires knowing what the agent was supposed to do, or what it meant, it probably isn't a lint rule. If it only requires what the trace records (plus facts you declare about your tools), it can be checked deterministically.

A few examples along that line:

  • Arguments that violate the tool's JSON Schema: deterministic.
  • A call to a tool that doesn't exist: deterministic.
  • A result declared as failed that's fed into a side-effecting action: deterministic.
  • The same call repeatedly with no apparent progress: not necessarily a bug, since it might be polling or a legitimate retry. This gets surfaced for review.
  • The agent picked the wrong strategy: a question for an eval.

tracelint includes several other structural checks, but they all follow the same principle. It reports a defect only when the trace provides enough evidence to prove it.

The example, in Phoenix

tracelint reads the OpenInference spans you already collect, with no changes to the agent. From Phoenix:

Figure 1. The run in Phoenix. The run reports OK (1), but the pipeline's output says "result": "UNSTABLE" (2) with 3 failing tests (3), and the agent's next tool call is deploy (4).

Run tracelint on this trace as-is and it won't fail it, and that's correct: nothing in the trace says UNSTABLE means failure. That's your CI's convention. Two things only you know have to come from you:

  • UNSTABLE (like FAILURE and ABORTED) means the pipeline failed
  • deploy takes a real action (it changes production), so acting on a failed result there is a defect, not just something to review

They go in a small tools.json:

{
  "tools": {
    "run_release_pipeline": {
      "metadata": {
        "failure_when": {"pointer": "/result", "in": ["FAILURE", "UNSTABLE", "ABORTED"]}
      }
    },
    "deploy": {
      "metadata": {"side_effecting": true}
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Run it again with that file:

$ tracelint check spans.json --format openinference --tools tools.json
689451f502ecfcb8bc0063e2615ed8c5: 2 finding(s), exit 2
  [hard_event] R2a tool_error_event  (step 3)
    'run_release_pipeline' returned a declared failure (/result='UNSTABLE')
  [hard_defect] R2b error_mishandled  (step 3,4)
    value(s) from the errored 'run_release_pipeline' result (b-4120-7f3a) reused
    as arguments to 'deploy' (a side-effecting action, no fallback)
  …
Enter fullscreen mode Exit fullscreen mode

In plain terms:

  • The pipeline failed. Its result was UNSTABLE, which your tools.json says counts as a failure.
  • Its build was deployed anyway. The build ID b-4120-7f3a from that failed run was passed to deploy, which changes production. That's a provable defect, so tracelint exits with code 2 and fails the CI job.

You don't have to write tools.json from scratch. tracelint init spans.json --format openinference -o tools.json drafts one from the trace, including tool schemas where the instrumentation records them.

Hard defects, candidates, and "not checked"

Every finding lands in one of three buckets, and the distinction matters more than any individual rule.

  • Hard defects are proven by the trace. In CI, tracelint check returns exit code 2 and fails the job.
  • Candidates are possible problems, such as repeated calls that might be a loop or might be a legitimate retry. They're shown for review but never fail the build. That's deliberate: one noisy heuristic is enough to make engineers disable a linter.
  • Not checked means a rule lacked the evidence it needed, and tracelint says so instead of counting it as a pass.

For example, the tools.json above doesn't say what a failed deploy looks like, so the same run's output includes:

suppressed (3) — not checked, not a clean pass:
  R2a tool_error_event: side-effecting tool 'deploy' returned an unclassifiable
  result and declares no failure_when predicate — cannot verify it did not fail
  …
Enter fullscreen mode Exit fullscreen mode

tracelint won't claim the deploy itself succeeded, because nothing tells it what failure would look like.

If you don't save traces in your tests yet, tracelint has a capture helper and a pytest fixture that record a run through your framework's OpenInference instrumentor, plus a GitHub Action for CI. The repo has setup details for each.

Where deterministic checks stop

  • tracelint's structural checks don't tell you whether the final answer is semantically correct. That's a separate, task-specific evaluation problem.
  • Some checks depend on facts you declare, such as which tools have side effects. If those declarations are wrong, the verdicts will be too.
  • Repetition and unexplained argument values stay candidates, because retries and value transformations can be legitimate.

How this fits with evals and production investigation

These layers overlap, but each is best suited to a different part of the reliability problem:

Layer Best suited for
Deterministic trace checks Explicit structural rules that can be proven from a single run
Evals Application-specific quality, behavior, grounding, tool choice, and policy
Production investigation Discovering recurring or previously unknown failure patterns across many traces

They also feed each other. Signal can surface a recurring trajectory problem in production. If that failure can then be expressed as a structural rule, a deterministic regression check can help keep it from coming back.

Try it on a Phoenix trace

pip install tracelint
tracelint check spans.json --format openinference
Enter fullscreen mode Exit fullscreen mode

The repo has setup details and a Phoenix guide. Feedback and issues are welcome, especially traces that tracelint doesn't handle well.

Top comments (0)