Have you ever shipped something an agent said it had finished, and only found out weeks later that nothing had actually run? I have. The failures that cost me time were never the crashes: a crash is a gift, because CI sees it. The expensive ones exit 0, print nothing unusual, and quietly skipped the work.
So instead of trusting exit codes I went looking for them in recorded traffic. Everything below is measured: one small agent task — fetch headlines, write a draft, verify it through a tool — running in three frameworks behind a single recording proxy that keeps every attempt. Aggregating that recording, the failures all wore one of three shapes:
- the verification never ran
- an invented argument got accepted
- the answer was empty
Code, traces and the analysis scripts are on GitHub:
sunnydachs
/
agent-framework-showdown
Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability
agent-framework-showdown
The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.
English | 日本語
Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.
The task
A tech-news digest agent:
- collect 5 headlines via a
fetch_headlinestool - write a ~100-word digest
- verify the word count via a
word_counttool, revising if out of band
All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.
How to run
# one venv per framework (Python 3.12 - CrewAI requires <3.14)
uv venv .venv-strands --python 3.12 && uv pip install --python .venv-strands/bin/python "strands-agents[litellm]"
uv venv .venv-langgraph --python 3.12 && uv pip install --python…Shape 1: the verification never ran, and the run still said success
I removed an argument from the verification tool so that it returned a static count instead of checking the draft. Nothing errored. The tool ran, answered, and the run reported success with a draft that was never verified.
- Strands: all 3 runs exited 0. Verifications actually executed: 0/3
- CrewAI: all 3 runs exited 0. Verifications actually executed: 0/3
- LangGraph: all 3 runs stopped with a hard error (its tool calls live in code)
LangGraph is the only one that cannot slip through silently: its structure cannot step over a broken tool.
There is a nastier variant, where the tool did answer with an error and the process still exited 0.
Strands, argument-type level:
tool error responses = 2, 1, 5 (three runs)
process exit code = 0, 0, 0
Those are not schema rejections. The call was well-formed; the tool answered with a processing error — "document 88 not found" — and every run still finished green.
Shape 2: the model invented a required argument, and the tool accepted it
I added a new required argument to the verification tool and never mentioned it in the prompt.
- Strands: all 3 runs had a plausible invented value accepted (0 errors)
- CrewAI: all 3 runs accepted (0 errors) — and all three invented the same sentence
- LangGraph: never reached the tool
The tool's only validation was "non-empty string".
Strands, 3 runs: "Check word count of AI agents news digest"
"AI agents digest word count check"
"Word count check for AI agents digest"
CrewAI, 3 runs: "Draft digest summarizing the five headlines into a flowing narrative"
Six fabricated values went straight through, and the recording is the only thing that shows where they came from.
That contrast is not purely a framework property, though: the two frameworks advertised the new parameter differently on the wire, and they did not sample the same way (see limitations).
Shape 3: the run said "done", and the answer was empty
This one only appeared in the model-driven frameworks.
Finish reason was stop, the visible answer was empty, and the draft was sitting inside the last tool-call argument.
- Strands: 2 runs (the content was in the tool argument: 100 and 106 words)
- LangGraph / CrewAI: 0 runs
A framework whose output node is the deliverable cannot produce this shape.
Where the evidence lives
what the process reports what actually happened
┌────────────────┐ ┌──────────────────────────────┐
│ exit 0 │ │ tool call #1 → error │
│ "done" banner │ ≠ │ tool call #2 → static stub │
└────────────────┘ │ verify → never ran │
└──────────────────────────────┘
evidence lives here → at the tool boundary (wire)
The error responses existed on the wire, and the tool layer even counted them. The process just never looked.
The evidence of a failure survives only outside the framework, at the tool boundary.
Invisible failures still cost money
A failure that gets reported as success keeps racking up wasted runs while nobody notices.
In a separate experiment (approval gates, crash recovery, duplicate execution, same proxy), it showed up in three ways:
- On crash recovery, the framework without state recording re-ran the whole task (2 LLM calls). The one with a checkpointer resumed with 0.
- An irreversible operation (in that experiment, publishing) executed twice under two different call IDs, and both returned success.
- Even the "did nothing" failure burned LLM calls: CrewAI 6, 5, 4; Strands 4, 5, 4.
Duplicates do not show up in the framework's own trace, either. Every duplicated call carried a different call ID, so on the wire they look like two unrelated requests.
And if the retry only rewords the same intent, an idempotency ledger keyed on the argument bytes cannot tell it apart:
call ID: X ──> ledger ──> publish executes ──> success
call ID: Y ──> ledger ──> publish executes ──> success
↑ keys differ, so it doesn't look like a duplicate
result: the side effect runs twice (both return success)
How to prevent it: the smallest check per situation
- Hobby pipeline → assert the deliverable is not empty
- In CI → assert the expected tool was actually called (never 0 times)
- Writing data → validate values (type, enum, known set)
- Handling approvals or billing → record approval and execution as separate states
- Need an audit → record at the boundary (the framework's own trace cannot reconstruct it)
The hobby case is the one I keep running into. Nothing errors, the log is calm, the artifact is empty or unverified — and you find out weeks later. If three lines of assertion catch that, I see no reason to leave them out.
The CI case is the same failure one layer up. If the only contract is the exit code, a skipped verification is invisible by construction.
"Did the verification tool run?" is a different question from "did the process succeed?"
For tools that write data, the invented argument is the dangerous one, because the value is well-formed. That no human ever wrote it is something only the wire recording can show.
And on approvals: approved once is not executed once. In these runs the same logical operation was settled twice under two different call IDs, and both returned success.
The two checks I actually use
Before anything downstream consumes the output, I assert exactly two things: the deliverable is not empty, and the tool I expected to run did actually run. Boring, a few lines each, and they would have caught every failure in this post.
Keys, ledgers and approval binding are for the moment the action touches money or another person. My honest read: most hobby pipelines do not need a ledger. They need the non-empty check and an exit signal they can trust.
Honest limitations
- 3 runs per cell. A trend, not a statistical claim.
- One model across all runs; a different model may reword or fabricate differently.
- The irreversible operation is simulated; whether a gate held is read from the recorded traffic.
- Fine instruction-following differences are hard to separate from the schema change itself.
- One framework sampled at temperature 0, and the parameter descriptions were not identical between the two — so the "varied vs repeated" contrast has at least two plausible causes.
This is a personal OSS project with no warranty. Use it at your own risk, and issues are welcome.
Top comments (1)
Both of your assertions live inside the process that did the work. That is also where they stop: once the consumer only holds the sender's account of a run, "the deliverable is non-empty" and "the expected tool was called" can both be satisfied by the account alone. I keep measuring a public append-only event log where independent agents append signed rows and an aggregator reads them later, and two of the shapes you drew show up there with no call IDs anywhere in the system.
On duplicates: of 19,926 rows, 600 sat in groups where author, event kind and body matched exactly, and every one of those rows carried a distinct id. So id-based dedup only removes what the read cursor redelivers at second boundaries. A sender re-appending the same claim goes straight through uncounted, which is your differing-call-ID diagram arriving by a completely unrelated route.
Timing is worse. There is a
runtime_msfield for how long the work took, and the read path never reads it. In 11 rows it was stamped earlier than the acceptance row that supposedly started that work. Schema-valid, so stored, and nothing objected. A signature settles who said a thing, never that anything ran. Your proxy is a clock the agent does not own, and that is what makes those two assertions carry weight. Past the point where the evidence becomes another sender's self-report, what would you take as the minimum assertion?