Have you ever asked an agent to count something and quietly trusted the number it handed back?
I did, until I stopped reading the answer and started reading the tool's own response. In one run the model answered 321 where the id list sitting in that same response held 322 matching ids. Nothing else about the run looked wrong.
- the same question through three frameworks, behind one recording proxy, one model throughout
- everything below is measured from the recorded traffic, not from the frameworks' own reports
sunnydachs
/
agent-framework-showdown
Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability
agent-framework-showdown
The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.
English | 日本語
Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.
The task
A tech-news digest agent:
- collect 5 headlines via a
fetch_headlinestool - write a ~100-word digest
- verify the word count via a
word_counttool, revising if out of band
All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.
How to run
# one venv per framework (Python 3.12 - CrewAI requires <3.14)
uv venv .venv-strands --python 3.12 && uv pip install --python .venv-strands/bin/python "strands-agents[litellm]"
uv venv .venv-langgraph --python 3.12 && uv pip install --python…The only thing I changed: what the tool hands back
Two versions of one tool. The first returns the id list and leaves the counting to the model. The second precomputes the count and hands that back.
A: the tool returns the rows B: the tool returns the count
┌───────────────────────┐ ┌────────────────────────┐
│ [3, 7, 11, … 330 ids] │ │ {count: 322, min, max} │
└───────────┬───────────┘ └───────────┬────────────┘
▼ the model counts ▼ the model reads
answered 321 (true 322) answered 322 (correct)
The grid is 3 frameworks x 2 tool modes x 3 list sizes x 7 threshold phrasings x 3 seeds, plus five extra 330-row probes — 408 recorded runs in total, every one behind the same recording proxy.
- mode
ids: the tool is asked for the id list only, the model must count - mode
stats: the tool is asked forcount/min/max - sizes 11, 110 and 330 ids, with phrasings such as
id >= 9,id > 9,no less than 9
408 runs in total, 407 of them exiting 0. A mode label describes what the tool was asked for, not what the framework carried into the prompt — so the analyzer also records what the last model call actually held: the id list, the count, or neither. That measurement is the reason the two columns below are not the two tool designs.
At 330 rows the rate matched three ways and the reasons only two
| Framework | model counts (ids) |
tool counts (stats) |
tokens/question ids
|
tokens/question stats
|
|---|---|---|---|---|
| CrewAI | 92.3% (24/26, 2 wrong) | 100% (26/26) | 14,600 | 12,565 |
| LangGraph | 92.3% (24/26, 2 incomplete) | 100% (26/26) | 8,574 | 143 |
| Strands | 92.3% (24/26, 2 wrong) | 100% (26/26) | 11,528 | 1,418 |
92.3% is the same number three times, and it is the same failure twice.
- CrewAI and Strands each answered two runs one or two off.
- LangGraph's two missing runs answered nothing at all: both generate to the completion ceiling and stop.
Three different tool-calling architectures — one keeps the tool call inside code-managed state, the other two let the model drive it — landed on the same rate. What they did not do was fail the same way, and the difference matters more than the rate.
At 11 and 110 rows the answer is right almost everywhere. The single exception is CrewAI in stats mode at 110 rows, which answered 106 where the true count was 107. That run had the tool's data in front of it, and it had also called the counting tool.
The wrong answers were not one kind of mistake
Five runs answered with the wrong number. Each was re-counted from the id list in that run's own recorded tool response, not from the analyzer's arithmetic:
| Run | Answered | True | Tool calls |
|---|---|---|---|
Strands, ids, id >= 9
|
321 | 322 | 1 |
Strands, ids, no less than 9
|
321 | 322 | 1 |
CrewAI, ids, id > 9
|
320 | 321 | 1 |
CrewAI, ids, id >= 9
|
322 | 323 | 1 |
CrewAI, stats, id > 9
|
106 | 107 | 1 |
The drift is always one or two, never an invented number. But it is not one clean arithmetic slip either: three of the five answered exactly the strict > 9 count to a >= 9 question, and the list held exactly one id equal to 9. A boundary read one operator too strict, or a tally one short — never a wild guess.
All five called the tool, and the tool's data was in the final prompt of all five — including the stats run whose correct count was sitting right there:
- the failure is in reading the data, not in fetching it
I found that out the hard way. My first version of this analysis said three of the five had skipped the tool call entirely — the counter only recognised some of the tool names the frameworks use, so CrewAI's calls were invisible to it. The failures are silent in the answer, and they were silent in my analysis too.
Only the frameworks that handed the count over got cheaper
| Framework | tokens/question, list reaches the model | tokens/question, count reaches the model |
|---|---|---|
| LangGraph | 8,574 | 143 (60x cheaper) |
| Strands | 11,528 | 1,418 (8x cheaper) |
| CrewAI | 14,600 | 12,565 (1.2x — and it never saw the count) |
LangGraph's stats runs are 60x cheaper because the count is computed in code and the model is handed a number instead of a list. Strands does the same and saves 8x. CrewAI is the exception on both axes: in its stats runs the tool's count never reached the model — the analyst was handed the id list again and counted it a second time, so the cost barely moved and the answer came from a list either way.
Read CrewAI's ids and stats columns as two samples of one task (24/26 and 26/26), not as a comparison between the two tool designs. A two-run gap at 26 runs is not a design effect.
What that says about the framework layer
the model the tool
┌──────────────────────┐ ┌─────────────────────────┐
│ reads the id list │ ───▶ │ returns 330 ids │
│ counts them itself │ └─────────────────────────┘
│ answers 321 │ ← the true count was here all along
└──────────────────────┘
no error · no retry · nothing to route differently
No error was raised in any of these runs. The tool was called, the answer had the right shape, and the process exited 0.
In this grid, at 330 rows, with one model and one endpoint, no framework moved the number: the failure is not absorbed by the graph, the agent loop or the callback plumbing. What did move it was arranging for the model to be handed the number instead of the list — and in one of the three frameworks that arrangement never took effect, because the plumbing between the tool result and the final prompt dropped it.
So what: the smallest change that removes it
| If you... | The failure that bites | The smallest thing that catches it |
|---|---|---|
| ask an agent to count rows | off-by-one answers with no error | compute the count in the tool and return it |
| hand back a list anyway | the model recounts what you already know | return count / min / max alongside the rows |
| design a tool you cannot see the output of | the count is dropped in the hand-off | assert the count is in the final prompt, not just in the tool result |
| report counts to a person | a number that looks fine and is not | diff the answer against the tool's own response |
The third row is the one this experiment cost me. A tool design is a hypothesis about what reaches the model; only the recorded prompt tells you whether it did.
Honest limitations
- 21-26 runs per cell: a direction, not a statistical claim. CrewAI's
ids-vs-statsgap is exactly two runs. - One model across every run, so a different model may drift differently or not at all.
- The two incomplete runs are runaway generations, not scored failures; one trace holds two ceiling-length calls, the other one. Both are excluded from the token means.
- The endpoint's daily cap killed runs mid-grid; those 170 labels were re-run successfully, and the successful attempt is what gets scored. Nothing was dropped from the grid on those grounds.
- The token figures are means over the runs that produced an answer, taken from the recorded usage. They compare what each framework put in front of the model, not the frameworks themselves.
Reproduce it
All 408 runs, their traces and the analysis scripts are open, and every number in this post is logged with the file it comes from and a command that recomputes it:
https://github.com/sunnydachs/agent-framework-showdown
This is a personal OSS project, so there is no warranty. Use it at your own risk, and issues are welcome.
Top comments (2)
The boundary operator confusion matches what breaks when agents parse raw arrays. Handing three hundred integers to a model forces it to filter and tally simultaneously in one generation pass, so boundary checks like strict versus inclusive comparison slip easily. Moving aggregation into the tool signature cuts prompt bloat from thousands of tokens down to double digits, while turning an in-context loop into a deterministic query before the model ever touches the data.
@sunnydachs This is one of the most rigorous agent studies I’ve read recently. The distinction between "model counting" vs. "tool computing" is exactly the gap that breaks production agents—and it’s why hallucination rates spike at scale (330 rows → off-by-one errors).
Your data proves that determinism must live in the code, not the context window.
I’m Harun (12yo founder of HYNAWEB), building KODA—an AI coding mentor designed for beginners who don’t have your infrastructure budget but still deserve accuracy.
While your experiment focuses on framework orchestration, KODA focuses on pedagogical correctness. We enforce a similar principle via our Constitutional AI Core:
✅ Article 4 (Truthfulness): Never invent APIs/numbers. If the tool doesn’t return the count, KODA refuses to guess.
✅ Article 5 (Code Honesty): Don’t silently change requirements. If a boundary condition (
>=vs>) is ambiguous, ask for clarification rather than assuming.Your finding that LangGraph saves 60x tokens by handing back
{count: 322}instead of[3, 7, ...]is brilliant. It validates that context hygiene = cost control.Quick question for your expertise:
In your Strands/CrewAI tests, did you observe any correlation between prompt verbosity and boundary error rates? My hypothesis is that verbose prompts increase cognitive load, making operators like
>=more prone to misinterpretation.Would love to hear your take. Your work raises the bar for everyone building agentic systems. 🐯️