<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yugesh Jha</title>
    <description>The latest articles on DEV Community by Yugesh Jha (@yugesh_jha_4493f0f45525c1).</description>
    <link>https://dev.to/yugesh_jha_4493f0f45525c1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4159956%2F4e23cfe1-4cb1-4eec-a1d9-2058bea0ac5d.png</url>
      <title>DEV Community: Yugesh Jha</title>
      <link>https://dev.to/yugesh_jha_4493f0f45525c1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yugesh_jha_4493f0f45525c1"/>
    <language>en</language>
    <item>
      <title>When an AI agent says it's done and it isn't</title>
      <dc:creator>Yugesh Jha</dc:creator>
      <pubDate>Sat, 03 Oct 2026 16:11:19 +0000</pubDate>
      <link>https://dev.to/yugesh_jha_4493f0f45525c1/when-an-ai-agent-says-its-done-and-it-isnt-5207</link>
      <guid>https://dev.to/yugesh_jha_4493f0f45525c1/when-an-ai-agent-says-its-done-and-it-isnt-5207</guid>
      <description>&lt;p&gt;&lt;em&gt;The edit is usually fine. The claim about the edit is the problem, and it is a harder one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You ask for a change across four files. The agent works for a few minutes, edits them, and says the change is complete. It is not. Nothing was run, one caller in a fifth file no longer compiles, and you find out in CI or in review or from somebody else.&lt;/p&gt;

&lt;p&gt;The frustrating part is that the edit was usually reasonable. What was wrong was the sentence at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it happens
&lt;/h2&gt;

&lt;p&gt;A language model produces the most plausible continuation of what came before. After a sequence of edits, the most plausible continuation is a summary saying the work is finished, because that is how the thousands of examples it learned from ended. Nothing in that process checks.&lt;/p&gt;

&lt;p&gt;So "done" is not a claim the model is making about your repository. It is the shape of a closing paragraph. The only thing that turns it into a real claim is a tool that actually ran something, and then reported what came back.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and why it is worse than a bad edit
&lt;/h2&gt;

&lt;p&gt;A wrong edit you catch in the diff costs a minute. A wrong claim costs you the assumption you were reviewing under.&lt;/p&gt;

&lt;p&gt;This compounds. The reason agents are useful is that you stop reading every line and start reading the summary. Once the summary is unreliable you have to go back to reading every line, and at that point the agent has moved the work rather than done it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to tell the difference
&lt;/h2&gt;

&lt;p&gt;The question to ask of any tool is simple: &lt;strong&gt;did it run anything, and can you see what came back?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Did it execute your tests, or describe executing them?&lt;/strong&gt; Those look similar in a summary and are not the same event.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the raw output on screen?&lt;/strong&gt; A tool that ran your suite has output. A tool that did not will paraphrase.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What does it do when it cannot verify?&lt;/strong&gt; The honest behaviour is to say so. A tool without that check tends to report success anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did it see the file it broke?&lt;/strong&gt; An agent reading only your open tabs cannot verify a change it cannot observe.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A test you can run in ten minutes
&lt;/h2&gt;

&lt;p&gt;Take a repository you know well with a test suite that passes. Ask the agent for a change touching at least three files. Then, before you look at the diff, break something it just wrote by hand and ask it to continue.&lt;/p&gt;

&lt;p&gt;A tool that runs your suite notices. A tool that does not will keep going and tell you everything is fine. That single exercise tells you more than any comparison page, including ours.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do if your current tool does this
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Make the tests the specification.&lt;/strong&gt; An agent iterating against a suite is only as good as the suite. Break the behaviour a test guards and confirm the test goes red, because a test that passes either way actively steers an agent wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shrink the unit of change.&lt;/strong&gt; Review fatigue is what turns an unverified claim into a merged defect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask for the commands.&lt;/strong&gt; If the tool can show what it ran, make it. If it cannot, treat every summary as a draft.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where AstraCode fits
&lt;/h2&gt;

&lt;p&gt;This is the problem we built &lt;a href="https://astracode.io" rel="noopener noreferrer"&gt;AstraCode&lt;/a&gt; around. It plans the change, makes it, runs your tests, reads the output, fixes what it broke, and runs them again, with what it ran on screen rather than summarised. When it cannot verify something it says so instead of reporting success. There is a free tier with no card, and the ten-minute exercise above is a fair way to judge it, or any other tool.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://astracode.io/blog/when-an-ai-agent-says-its-done-and-it-isnt/" rel="noopener noreferrer"&gt;AstraCode blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Cost per token is the wrong number for coding agents. Here's what we measured instead.</title>
      <dc:creator>Yugesh Jha</dc:creator>
      <pubDate>Sat, 03 Oct 2026 14:38:05 +0000</pubDate>
      <link>https://dev.to/yugesh_jha_4493f0f45525c1/cost-per-token-is-the-wrong-number-for-coding-agents-heres-what-we-measured-instead-1fp5</link>
      <guid>https://dev.to/yugesh_jha_4493f0f45525c1/cost-per-token-is-the-wrong-number-for-coding-agents-heres-what-we-measured-instead-1fp5</guid>
      <description>&lt;p&gt;We build AstraCode, an AI code editor whose agent has to prove its own work. Model calls are the largest line on our bill, so in September we stopped guessing and measured: eleven models, the same set of real coding tasks, each run more than once, scored on pass rate and on &lt;strong&gt;cost per finished task&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step count beats token price
&lt;/h2&gt;

&lt;p&gt;The cheapest model per token is not the cheapest model per task. On one task a model took 130 agent steps; another finished in 32. Every step re-sends context, so the step count multiplies everything else.&lt;/p&gt;

&lt;p&gt;It is not as simple as "expensive models wander more" either. Some of the priciest models took &lt;em&gt;fewer&lt;/em&gt; steps than the cheap one. We wrote that rule down, then had to cross it out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; log steps per task alongside tokens. A model that is 3x cheaper per token and takes 4x the steps is a more expensive model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing by prompt doesn't work
&lt;/h2&gt;

&lt;p&gt;The obvious optimisation is a router: read the request, send easy ones to a small model and hard ones to a big one. We tried three versions of it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A per-step router came out 0.4% &lt;em&gt;worse&lt;/em&gt; than not routing at all.&lt;/li&gt;
&lt;li&gt;A classifier that predicted difficulty from the prompt reached a rank correlation of about 0.75 and still lost.&lt;/li&gt;
&lt;li&gt;Even an oracle with perfect hindsight picked the small model on every task, so there was no headroom to win.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reason is variance. The same small model passed one task &lt;strong&gt;67% of the time on the identical prompt&lt;/strong&gt;. Whether a run succeeds is decided during the attempt, not by the prompt, so no prompt classifier can see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detection does work
&lt;/h2&gt;

&lt;p&gt;What does work is checking the result:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run the small model.&lt;/li&gt;
&lt;li&gt;Check its work. Run the tests, then undo the change and confirm those tests fail without it, so a test that passes either way doesn't count as proof.&lt;/li&gt;
&lt;li&gt;Only if a check fails, hand the task to a stronger model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Pass rate went from &lt;strong&gt;0.64 to about 1.00&lt;/strong&gt;, at roughly &lt;strong&gt;$0.20 per task&lt;/strong&gt; against $0.03 for the small model alone.&lt;/p&gt;

&lt;p&gt;These are our tasks and our harness, so treat them as one data point, not a ranking. Every model we tested is good at something.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we changed
&lt;/h2&gt;

&lt;p&gt;AstraCode no longer asks you to pick a model. AstraOne starts with the cheaper model and escalates when a check fails. The checks are the same ones you'd want anyway: tests that are shown to fail without the change, a diff you review hunk by hunk, and a checkpoint before every turn.&lt;/p&gt;

&lt;p&gt;If you're building on LLMs, two habits are worth more than any model choice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;measure &lt;strong&gt;cost per finished task&lt;/strong&gt;, not per token;&lt;/li&gt;
&lt;li&gt;run every task &lt;strong&gt;several times&lt;/strong&gt; before you trust a number.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AstraCode is free to start, no card: &lt;a href="https://astracode.io" rel="noopener noreferrer"&gt;https://astracode.io&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>coding</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
