<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rashid Mahmood</title>
    <description>The latest articles on DEV Community by Rashid Mahmood (@code-with-rashid).</description>
    <link>https://dev.to/code-with-rashid</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1268469%2Fdca06eca-ff9e-476c-be84-206a0aa9cc9e.png</url>
      <title>DEV Community: Rashid Mahmood</title>
      <link>https://dev.to/code-with-rashid</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/code-with-rashid"/>
    <language>en</language>
    <item>
      <title>What agent frameworks cost on the wire: measurements from agentic-arena</title>
      <dc:creator>Rashid Mahmood</dc:creator>
      <pubDate>Wed, 30 Sep 2026 17:05:46 +0000</pubDate>
      <link>https://dev.to/code-with-rashid/what-agent-frameworks-cost-on-the-wire-measurements-from-agentic-arena-13o7</link>
      <guid>https://dev.to/code-with-rashid/what-agent-frameworks-cost-on-the-wire-measurements-from-agentic-arena-13o7</guid>
      <description>&lt;p&gt;Most framework comparisons argue from feature lists. I wanted numbers, so &lt;a href="https://github.com/code-with-rashid/agentic-arena" rel="noopener noreferrer"&gt;agentic-arena&lt;/a&gt; holds the model, tools, datasets and iteration budget fixed and measures what each framework does differently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read this first:&lt;/strong&gt; everything below is measured against a scripted (mock) model, so the turns are byte-identical for every framework. That makes these wire and behaviour measurements. They say nothing about which framework writes better answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What a framework puts on the wire
&lt;/h2&gt;

&lt;p&gt;Mean prompt tokens per item on a 15-item tool-use task:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;framework&lt;/th&gt;
&lt;th&gt;prompt tokens&lt;/th&gt;
&lt;th&gt;vs baseline&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;vanilla (stdlib loop)&lt;/td&gt;
&lt;td&gt;753.5&lt;/td&gt;
&lt;td&gt;1.00x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;langgraph&lt;/td&gt;
&lt;td&gt;753.5&lt;/td&gt;
&lt;td&gt;1.00x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pydantic_ai&lt;/td&gt;
&lt;td&gt;794.0&lt;/td&gt;
&lt;td&gt;1.05x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;microsoft_af&lt;/td&gt;
&lt;td&gt;802.0&lt;/td&gt;
&lt;td&gt;1.06x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;google_adk&lt;/td&gt;
&lt;td&gt;836.1&lt;/td&gt;
&lt;td&gt;1.11x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;openai_agents&lt;/td&gt;
&lt;td&gt;856.9&lt;/td&gt;
&lt;td&gt;1.14x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;smolagents&lt;/td&gt;
&lt;td&gt;2935.5&lt;/td&gt;
&lt;td&gt;3.90x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Six of seven sit inside a 1.15x band, and none is leaner than the hand-rolled loop. smolagents' 3.90x is a 4,207-character system prompt where the arena asked for 384, including a prose restatement of tools it already sent as a schema. The tools are transmitted, and billed, twice.&lt;/p&gt;

&lt;p&gt;That is a worst case. The same conversation run to 30 turns shrinks the gap from 8.83x on the first request to 1.27x on the 31st, because a fixed overhead decays as the conversation grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What breaks when the provider fails
&lt;/h2&gt;

&lt;p&gt;I scripted 429, 500 and 400 responses. The hand-rolled baseline has no retry at all, so one 429 loses the item. Every framework survives a single 429. Only smolagents survives three in a row, by quietly sleeping roughly two to four minutes on one item. The item passes, so nothing in a scorecard shows it, but in a batch your throughput quietly collapses.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What delegation costs
&lt;/h2&gt;

&lt;p&gt;A three-role pipeline doubles the LLM calls and multiplies prompt tokens by 2.5x, whether you build it with a graph library or a or loop. The structure costs that, not the framework.&lt;/p&gt;

&lt;p&gt;Model-decided handoffs cost about 10% more prompt than the same pipeline wired structurally. Of that gap, 94% is the    ransfer_to_* tool schemas, which ride on every request whether or not anyone delegates. You pay for the options you offer, not the ones you use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark found bugs in itself
&lt;/h2&gt;

&lt;p&gt;Five of the seven adapters never passed the shared request timeout to their client, so they waited out their library's default through a 20-second hang on a one-second budget. A hanging mock provider exposed it. It is fixed, and gated in CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Everything above has a command that regenerates it, and CI regenerates each number on a clean install:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/code-with-rashid/agentic-arena
&lt;span class="nb"&gt;cd &lt;/span&gt;agentic-arena
python &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; arena run &lt;span class="nt"&gt;--arena&lt;/span&gt; tool_use &lt;span class="nt"&gt;--framework&lt;/span&gt; all &lt;span class="nt"&gt;--mode&lt;/span&gt; mock &lt;span class="nt"&gt;--no-scorecard&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Full findings and the reasoning behind each: &lt;a href="https://code-with-rashid.github.io/agentic-arena/findings/" rel="noopener noreferrer"&gt;https://code-with-rashid.github.io/agentic-arena/findings/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
