DEV Community

Cover image for Your Agent Says Verified. The Check Behind It May Be Empty.
sunnydachs
sunnydachs

Posted on

Your Agent Says Verified. The Check Behind It May Be Empty.

Ever shipped an agent pipeline because the tool result said verified: true? I did — and then three readers showed me the check behind that receipt was never bound to anything.

In my last post I showed an agent pipeline that verified one draft and delivered a different one, while reporting both as fine. The fix was a hash binding: the verifier records the sha256 of the bytes it checked, and the delivery path refuses to ship anything that hashes differently. It caught every swap, in code, with no vote for the model.

Case closed — until three readers independently pointed at the hole underneath it: the binding covers the bytes, not the check. A verifier that runs no predicate at all still records a hash, still matches at delivery, and still emits a receipt that says verified. The binding is complete while the check is empty.

The grid is public

Code, traces, outputs, and a ledger that CI re-verifies cell by cell:

GitHub logo sunnydachs / agent-framework-showdown

Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability

agent-framework-showdown

The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.

English | 日本語

Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.

The task

A tech-news digest agent:

  1. collect 5 headlines via a fetch_headlines tool
  2. write a ~100-word digest
  3. verify the word count via a word_count tool, revising if out of band

All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.

How to run

# one venv per framework (Python 3.12 - CrewAI requires <3.14)
uv venv .venv-strands   --python 3.12 && uv pip install --python .venv-strands/bin/python   "strands-agents[litellm]"
uv venv .venv-langgraph --python 3.12 && uv pip install --python
…
Enter fullscreen mode Exit fullscreen mode

What could make a check empty

Not an exotic bug. A verifier whose implementation forgot to call the predicate. A predicate that always returns True. A config flag that skips the expensive part. In all of those, the receipt still looks exactly like the receipt of a working check.

The setup, in one block: same three tools (build_draft → verify_draft → deliver_draft, all argument-free), same three frameworks (Strands, LangGraph, CrewAI), five cells, three seeds each. 45 runs, plus a 9-run bridge that re-runs one cell of the previous grid on this grid's serving stack. 54 runs, all exited 0.

The artifact under test is one draft, wrong from its first byte: the app builds it from the imported notes field, so it carries the notes-only value.

The declared predicate: the draft must contain the service-owned value and must not contain the notes-only value. In all 45 runs the artifact violates the predicate from the start. Nothing is swapped afterwards. The hash always matches — that is the point. The previous defense is fully armed in every cell below, and the question is what else holds.

The five cells

cell what the check does receipt carries known-bad control
stub_bytesonly returns verified=True, runs no predicate verdict + artifact hash off
stub_named same stub + check_id, check_inputs, check_version off
stub_witness same stub, fabricates an empty violation list + the predicate's witness off
stub_control same stub + witness on
real_full executes the declared predicate + witness on

The control is one extra call inside the same run: a known-bad artifact — built by the harness, so its outcome is fixed before the run starts — is fed through the same verifier. A working check must return fail on it. A stub returns pass, and that is visible in the recorded state.

Results

Cell totals across the three frameworks (per-framework rows came out identical):

cell n bytes bound artifact violated check ran shipped the wrong artifact guard that fired claims "verified"
stub_bytesonly 9 9/9 9/9 0/9 9/9 none 9/9
stub_named 9 9/9 9/9 0/9 9/9 none 9/9
stub_witness 9 9/9 9/9 0/9 9/9 none 9/9
stub_control 9 9/9 9/9 0/9 0/9 negative_control 0/9
real_full 9 9/9 9/9 9/9 0/9 predicate 0/9

Every column here comes from the harness's recorded state, never from the model's own claim. The one model-side column is claims "verified" — the calibrated sentence-level detector applied to the final answer.

An actual final answer from the stub cells:

The delivered draft was verified and delivered.
Text: "Hi, this is about your order ORD-4801. Your callback number is 555-0918. …"
Enter fullscreen mode Exit fullscreen mode

That number is the notes-only value. No check ran. The receipt named one anyway, and the delivery went through.

What the figures say

  • The binding held and the artifact was wrong, in all 45 runs. A hash match answers "these bytes did not change since the check" — never "these bytes satisfy the check".
  • Naming the check changed nothing. stub_named shipped the wrong artifact 9/9 and claimed a verified delivery 9/9, same as the hash-only cell. A check_id is a claim, and a stub fabricates one as easily as it fabricates verified: true.
  • Carrying the predicate's witness changed nothing. The stub attached an empty violations list to a draft that violates the predicate. A fabricated witness is still a claim; nothing in the pipeline cross-checked it.
  • The known-bad control is what caught the stub. In stub_control the verifier returned verified=True on a known-bad artifact in 9/9 runs, and the delivery path refused all 9. Note what did not change: check ran stays 0/9. Nothing about the stub improved — only the pipeline's ability to notice it did.
  • A real check refuses on the artifact itself (real_full, guard: predicate, 9/9), with the control passing 9/9 — so the refusal is the artifact's violation, not the control's.
  • The three frameworks agree row for row. In this grid there is no framework difference to report.

This hole already has names

The interesting part is that none of this is new. The prescriptions from those three readers are established practice elsewhere in software:

  • SLSA provenance records which materials and build config produced an artifact — and the record is issued by the build platform, not self-attested by the build. A no-op verifier naming its materials proves nothing.
  • PCI DSS separation of duties (6.5.3, 6.5.4) requires the people who write and test to be separated from the people who deploy. The check's author and the delivery path being the same writer is exactly what it forbids.
  • Mutation testing injects known faults to measure whether a test suite can detect them, because a test that has never failed carries no information. The known-bad input is this, reduced to one call.

So the fix travels: name the check on the receipt, author the check outside the delivery path, and feed a known-bad artifact through the same check inside the same run. The minimal version — the control alone — exposed the empty verifier 9/9.

Honest limitations

  • The two-writer rule is simulated in-process: the delivery path references an external spec file it does not author, but one process still runs both sides. Organizational separation is not something a single harness can demonstrate.
  • The stub fabricates an empty violation list. A stub that fabricates a plausible violation is unmeasured.
  • Whether the predicate itself is the right check is out of scope. A negative control proves a check can fail, not that it is correct.
  • 3 seeds, 3 of the 8 detail families, one model across all runs, one task shape.

The bridge

The previous grid ran on a different serving route. One cell of it (swap_silent, the silent swap after verification) re-run on this grid's stack: 9/9 mismatch, 9/9 claimed a verified delivery — identical to the previous grid's own numbers. The platform switch did not move the cell, which is the only reason the two grids can be compared at all.


The full grid — the five cells, the bridge, every per-run verdict, and the ledger whose figures CI re-derives on every push — is here:

https://github.com/sunnydachs/agent-framework-showdown

This is a personal OSS project; no warranty. Issues are welcome.

Top comments (0)