What happens when enterprise requirements - human approval gates, audit
trails, structured output - hit three agent frameworks? The first article
m...
For further actions, you may consider blocking this person and/or reporting abuse
Using one recorder proxy across all three frameworks makes this comparison much more informative than feature checklists. The duplicate publish call in Strands is especially important: an approval gate can hold while the side effect is still unsafe unless the execution layer deduplicates it. I’d be interested in a follow-up that scores approval, idempotency, and audit reconstruction together with the same operation ID. For LangGraph, was the missing tool order/argument data absent from the raw provider traffic too, or only from the framework-level trace surface?
Thanks — and Great question — I went back to the raw provider traffic to answer it properly: absent from the raw traffic too, not just the trace surface. Across all 21 LangGraph trace files, zero tool_call IDs appear anywhere in the recorded wire traffic, and the LLM requests carry plain-text prompts with no tools field (vs 53 records with tool_call IDs in Strands). In this setup the graph's tools are plain Python functions called in code — the model never sees a tool schema on the wire. One caveat worth adding: a LangGraph setup using bind_tools + ToolNode would put tool calls on the wire; my graph routes through code deliberately, which is exactly the audit trade-off the post describes.
The operation-ID scoring idea (approval + idempotency + audit reconstruction in one score) is a strong follow-up — the double-fired publish in Strands already shows why it matters: two different call IDs, both returning success, no dedup at the execution layer.
Scoring approval, idempotency and audit on one operation ID is the right follow-up. An approval gate that allows once but cannot deny the replay still moves money twice.
If anyone has a redacted log of agent tool calls that should have been single-use, we will run a claim-before-execute rule over it in shadow mode and send back a CSV of allow, hold or deny with reason codes. Which shows up first in your traces: timeout retry, or the model emitting two tool calls with different IDs?
From the recorded traffic: the different-ID double-emit comes first — it needed no timeout at all, just the model firing twice after approval. The retry shape is real too: when I forced a transient-504 retry, the model reworded the args and a byte-based key missed the duplicate every time. Approve-once needs an operation ID that survives re-reasoning, not just replay. Thanks for the offer — the earlier runs' traces are public in the repo, so your rule can run over them as-is.
Agree on the ordering: different-ID double-emit after approval is the scarier failure, and a byte key that misses a reworded retry is the other. An operation ID that survives re-reasoning is the approve-once control we care about above money tools.
If the earlier runs' traces are still public in the repo, send the path and we can shadow-run a claim-before-execute rule across those attempts and return would-have allow, hold, and deny counts with reason codes. Nothing blocks production.
When approval is granted once, do you bind that operation ID to amount and destination before any provider call so a later re-plan cannot reuse the same yes on a different payload?
Still public — same repo,
github.com/sunnydachs/agent-framew...
retry traces under traces/ (e.g. traces/llm_calls_strands__strands__idem_position_reworded_run1.jsonl), ledger exports in traces/idem/, exec logs under outputs/ (e.g. outputs/idem_retry_strands__idem_position_reworded_run1.json)On binding: yes, at claim time. The PENDING row stores the parameter hash before the provider call, so the same key with a different payload is refused as a caller bug rather than reusing the yes. Measured in the reworded runs: 8 refusals, 0 re-executions, one publish per run; and in the same-args runs the retry returned the recorded result without executing again.
The finding that LangGraph's
interrupt()enforces the gate structurally while Strands leaves it to the model is the whole ballgame for regulated work. We run agents that can publish and delete, and the rule we converged on is: any destructive action has to route through an edge/state the model cannot talk its way past. Prompt-only gates fail exactly when you most need them — the run that's confident and wrong is the one that "asks" and then publishes anyway. The "empty output while exiting 0" from Strands is a great catch and lines up with what I see: model-driven loops fail silently, graph-driven ones fail loudly, and loud failures are cheaper. Two things I'd love to see in a follow-up: (1) resume latency after a crash mid-interrupt, not just a clean pause — the checkpointer restoring in 0.0s is impressive but the real test is process death; (2) whether the audit trail survives a retry without double-counting the approval. Nice methodology keeping the recorder proxy and model fixed across runs.Thanks — and your rule is exactly where I landed too: the gate has to live in the graph's structure, not in the model's compliance. Both follow-ups go on my list:
The double-firing publish in Strands is the failure that bites hardest once actions touch an external API. We ran into the exact same pattern: prompt instructions said to execute once after approval, but the model-driven loop emitted two consecutive tool calls with identical payloads. You can tell the model not to double-fire, but the only reliable guard is an idempotency key at the execution boundary, either derived from the draft hash or passed as a run-scoped token to the downstream API.
The empty output exiting 0 is just as sneaky. Runners checking only exit status treat a polite handover to a validator tool as a complete pass. We ended up gating the runner exit on non-empty payload length and tracked file hashes rather than exit code alone.
Thanks — the idempotency key at the execution boundary is the fix I'd converge on too. In my traces the double-fired publish returned success both times with two different tool_call IDs, so the tool itself has no guard — dedup has to live outside the model. Your payload-length gate for the empty-output failure is a nice cheap pattern; the trace holds the full result in the tool args, so a non-empty check would have caught all three empty runs.
The cost here is not the model — it is that every approval gate and audit-log write is another metered step, and frameworks that "fail differently" also bill differently. Once compliance forces human-in-loop plus retention, your cost-per-successful-task stops being a model question and becomes an orchestration one. The EU AI Act angle is the same shape as APAC data-residency rules: the regulation ends up deciding your architecture, not the benchmark. Did any framework audit-trail overhead change your routing decision, or was compliance the non-negotiable regardless of cost?
Good framing. In my runs the audit logging itself added no metered steps — the recorder sits at the proxy, outside the model path — so the cost asymmetry lived in the failure modes: a rejection re-ran the task in CrewAI (2 LLM calls) while LangGraph's checkpointer resumed with zero. On your question: compliance was the filter, cost ranked the survivors — the model ended up the cheapest line item.
The recorder sitting outside the model path is exactly why flat per-call billing hides this — the meter that matters is the orchestration layer, not the model. Your two-stage read (compliance filters, cost ranks survivors) is the cleanest framing I've seen for why routing can't be a single score. Did cost ever actually override the compliance gate in your runs, or was compliance always hard-gated and cost only re-ordered what survived?
Thanks! Always hard-gated — cost never overrode compliance in these runs. To be honest, the benchmark never faced that temptation: every framework ran every cell, so the two-stage read is what the data implies for a routing decision, not a trade I actually had to make.
The silent empty output is the one that'd keep me up at night. We hit something similar last year - stage exited clean, returned nothing, and the consumer just quietly processed an empty result for two hours before anyone noticed. I'd want that trace front and center before any other dashboard, not buried behind a second tab you only open when something explodes. The LangGraph audit note is worth pairing with this since you can end up with good evidence split across the wire and the code base, which is its own recovery problem.
Two hours of quiet empty-processing is exactly why I record every call — though the data was never lost: the full result was sitting in the last tool-call argument in all three runs.
The scariest cell in this write-up isn’t a loud crash — it’s success-shaped emptiness: exit 0, empty final payload, full result only living in a tool arg.
That pairs with the gate finding: prompt-only approval is compliance theater; a destructive action has to route through structure the model can’t talk past (graph edge / interrupt), and the execution boundary needs idempotency so “approved once” can’t double-fire.
I’d treat non-empty deliverable + operation-id dedup as first-class eval contracts — same class as “did the test actually run?” — before ranking frameworks on vibes.
"Success-shaped emptiness" — the phrase I wish I'd had in the article. Non-empty deliverable as a first-class eval contract: agreed.
This is the kind of benchmark I find much more useful than a simple “which framework is faster?” comparison. The interesting part is seeing how each framework fails when you add the messy requirements that show up in real systems—approval gates, auditability, retries, and structured output.
The “success-shaped empty result” really stood out to me. A system returning exit 0 while quietly producing nothing is exactly the sort of issue that can survive testing and become painful in production.
I also appreciate the honesty around the limitations—one model and a relatively small number of runs makes this directional rather than definitive. But that actually makes the methodology feel more trustworthy.
The bigger takeaway for me is that choosing an agent framework isn't just about developer experience or features anymore. You have to think about where control lives, where evidence lives, and what happens when things go slightly wrong. Really interesting work—I'd definitely be curious to see the crash/retry and idempotency experiments you hinted at next.
Thanks — glad the limitations section read that way.
Good news: the crash/retry + idempotency experiments became the follow-up article (34 runs):
dev.to/sunnydachs/do-agents-surviv...
Two headlines from it: a durable checkpointer resumed a SIGKILLed process in 0.01s with zero extra LLM calls (the others re-ran: 5.3 and 2.0 avg calls), and a content-hash idempotency key failed silently when the model reworded the args on retry — the duplicate side effect slipped through, exit 0.