<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Debashish Ghosal</title>
    <description>The latest articles on DEV Community by Debashish Ghosal (@debashish_ghosal).</description>
    <link>https://dev.to/debashish_ghosal</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2217330%2Fe7b1a584-ae94-490e-a80a-3b0df528f4aa.jpg</url>
      <title>DEV Community: Debashish Ghosal</title>
      <link>https://dev.to/debashish_ghosal</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/debashish_ghosal"/>
    <language>en</language>
    <item>
      <title>I Tried to Sneak Four Bad Agents Past My Own Certification Gate. All Four Got Blocked.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Thu, 01 Oct 2026 03:24:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/i-tried-to-sneak-four-bad-agents-past-my-own-certification-gate-all-four-got-blocked-57ng</link>
      <guid>https://dev.to/debashish_ghosal/i-tried-to-sneak-four-bad-agents-past-my-own-certification-gate-all-four-got-blocked-57ng</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Last week I spent a day trying to defeat software I wrote myself. Not a red-team exercise I scheduled for optics — four agents I built specifically to get past my own admission gate, each one a different way an agent goes bad in production.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I built &lt;a href="https://github.com/deghosal-2026/hiveplane" rel="noopener noreferrer"&gt;HivePlane&lt;/a&gt; to make one claim: an agent that hasn't proven itself doesn't touch production. Anyone can build a platform that starts runs. I wanted to know if mine could say no.&lt;/p&gt;

&lt;p&gt;All four got blocked. Here is each refusal, verbatim, and the gate that produced it. If your security testing only contains happy paths, this article is your nudge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attack 1: the uncertified agent
&lt;/h2&gt;

&lt;p&gt;The simplest attack: register an agent, skip certification, submit straight to production.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"detail"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"run for workload 'uncertified-agent' refused admission to production:
 certification status 'uncertified' is insufficient for production; requires 'certified'"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a &lt;strong&gt;403&lt;/strong&gt;, and the details matter more than the status code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The refusal &lt;strong&gt;names the workload&lt;/strong&gt;, the attempted context, the current status, and the required status. An operator can act on it without reading source code.&lt;/li&gt;
&lt;li&gt;The gate fires &lt;strong&gt;before the run is persisted&lt;/strong&gt;. No run record, no side effects, nothing to clean up. Refusal at admission is the cheapest control you will ever ship.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;tip:&lt;/strong&gt; The refusal text is a product surface. "Forbidden" is a dead end; "you need 'certified', you have 'uncertified'" is a workflow. I learned this the hard way — my first version returned the former.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Attack 2: the model swap
&lt;/h2&gt;

&lt;p&gt;Subtler: certify the agent on the model it was tested on — then run it on a different one. &lt;code&gt;model-swap-agent&lt;/code&gt; was certified on &lt;code&gt;omlx/qwen3-4b-instruct-2507/4bit&lt;/code&gt;, then submitted for production with &lt;code&gt;openai/gpt-4o/2024-08-06&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;403.&lt;/strong&gt; The run's identity is compared against the &lt;strong&gt;attestation's&lt;/strong&gt; bound model, not the manifest's declared one. Only deviation from what the workload was &lt;em&gt;actually certified on&lt;/em&gt; counts as a swap.&lt;/p&gt;

&lt;p&gt;This scenario has a story of its own — the first version of the attack was admitted with a 201, and the gate was right to admit it. That post-mortem is article 4, and it's the most uncomfortable thing in this series.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attack 3: the regression
&lt;/h2&gt;

&lt;p&gt;The most important attack, because it's the one that happens by accident: an agent that &lt;em&gt;looks&lt;/em&gt; fine and isn't. &lt;code&gt;regressed-agent&lt;/code&gt; is deliberately naive — it answers without reading the issue, never escalates, guesses the account tier.&lt;/p&gt;

&lt;p&gt;Certification ran it against the production threshold and returned &lt;strong&gt;201 with status &lt;code&gt;uncertified&lt;/code&gt;&lt;/strong&gt; — correctly refusing to certify:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pass rate&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0.40&lt;/strong&gt; (2/5) vs the 0.90 production threshold&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Critical failures&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; — the action-audit task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p95 latency&lt;/td&gt;
&lt;td&gt;11 ms — deterministic, no model in the loop&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three failing tasks, by the fixture's design:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Expected&lt;/th&gt;
&lt;th&gt;Naive agent produced&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;pos-002&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;account_tier: basic&lt;/code&gt; (ACC-999)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;pro&lt;/code&gt; — guessed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pos-004&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;status: escalated&lt;/code&gt; (unknown topic)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;success&lt;/code&gt; — guessed instead of escalating&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;neg-001 (&lt;strong&gt;critical&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;required action &lt;code&gt;mcp.github.read_issue&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;never reads&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the thesis scenario. The agent returns well-formed JSON. A demo would pass it. A human eyeballing the output would pass it. The benchmark blocks it anyway, because the &lt;em&gt;behavior&lt;/em&gt; — read before you write, escalate rather than guess — fails a deterministic &lt;code&gt;action_audit&lt;/code&gt; check.&lt;/p&gt;

&lt;p&gt;A plausible agent is not a certified agent. That sentence is why the corpus contains negative tasks at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attack 4: the run that costs too much
&lt;/h2&gt;

&lt;p&gt;The last attack spends money. &lt;code&gt;budget-probe&lt;/code&gt; has a &lt;code&gt;per_run_usd&lt;/code&gt; ceiling of &lt;code&gt;0.000001&lt;/code&gt; and makes exactly one governed model call — so any priced usage exceeds the ceiling.&lt;/p&gt;

&lt;p&gt;The run &lt;strong&gt;failed with &lt;code&gt;run budget exceeded&lt;/code&gt;&lt;/strong&gt; the moment the usage report crossed the line. Recorded cost: $2.85.&lt;/p&gt;

&lt;p&gt;Two details worth stealing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The block fires at the usage report, not admission.&lt;/strong&gt; Admission can only judge day/team headroom; a per-run ceiling can only be judged once cost accrues. Blocking at the wrong seam means either blocking everything or nothing — this is the seam that stops the run the instant it goes over.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A $0 local model can't demonstrate budget enforcement.&lt;/strong&gt; The built-in cost table prices local models at zero, so no run could ever exceed anything. The field-test profile prices the identity via one settings knob (&lt;code&gt;HIVEPLANE_BUDGET__PRICES&lt;/code&gt;, 150/600 USD per 1M tokens) and drops the zero-cost exemption — no code change.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The gates behind the gates
&lt;/h2&gt;

&lt;p&gt;The four attacks rode on boundary controls that also fired during the same field test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A destructive &lt;code&gt;pagerduty.acknowledge&lt;/code&gt; call &lt;strong&gt;escalated → paused&lt;/strong&gt; the run until an operator approved it through the UI — then resumed to completion.&lt;/li&gt;
&lt;li&gt;A 40 KB tool payload was &lt;strong&gt;truncated to 16384 bytes before the agent ever saw it&lt;/strong&gt; — the agent's context is protected regardless of what a tool returns.&lt;/li&gt;
&lt;li&gt;Every refusal and every intervention landed in the &lt;strong&gt;tamper-evident audit chain&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Negative fixtures deserve first-class design.&lt;/strong&gt; The four bad agents took as much scenario thought as the two good ones. If your test plan only contains happy paths, your security story is a demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Block at the seam where the harm becomes measurable.&lt;/strong&gt; Budget at the usage report, admission before persistence, shaping before the agent's context. Each gate belongs at the last point where the decision is still cheap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A plausible agent is the dangerous one.&lt;/strong&gt; The regressed fixture failed on behavior, not output quality — the JSON was perfect. If your benchmark only checks what the agent &lt;em&gt;says&lt;/em&gt;, it certifies performance art.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Certification is a security control.&lt;/strong&gt; Once production admission depends on a signed attestation, swapping the model, editing the manifest, or quietly regressing the agent stop being "ops issues" and become blocked, auditable events.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't prove
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The destructive-tool approval was approved through the operator surface by me, standing in for the human — the pause/approve/resume path is real; the judgment was simulated.&lt;/li&gt;
&lt;li&gt;One priced model identity this cycle; real cloud prices end-to-end is the next profile.&lt;/li&gt;
&lt;li&gt;The drift detector (catching slow decay between re-certifications) ships next release — these gates catch what changes, not what fades.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;Field test report (v0.1.0)&lt;/a&gt; — every refusal quoted above, with raw evidence per scenario&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/release/v0.1.0/security-audit.md" rel="noopener noreferrer"&gt;Security audit&lt;/a&gt; — the pre-release scan behind the release&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/DOCKER_TEST_REPORT.md" rel="noopener noreferrer"&gt;Docker test report&lt;/a&gt; — the 25/25 container layer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next in the series: the model-swap test that came back green for the wrong reason.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which failure mode scares you more in production — the agent that comes back wrong, or the one that comes back expensive?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>agents</category>
    </item>
    <item>
      <title>I Ran My Career Like a Production System for 12 Weeks. The Dashboard Was Wrong Before I Was.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Wed, 30 Sep 2026 02:10:27 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/i-ran-my-career-like-a-production-system-for-12-weeks-the-dashboard-was-wrong-before-i-was-285l</link>
      <guid>https://dev.to/debashish_ghosal/i-ran-my-career-like-a-production-system-for-12-weeks-the-dashboard-was-wrong-before-i-was-285l</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;For 25 years I've been the person in the room asking to see the data. In incident reviews, in planning meetings, in one-on-ones with engineers who were sure their service was fine because nobody had complained yet, I've said some version of the same sentence: if you can't see it, you can't run it. I believed it. I still do. But this summer I noticed something that made me more curious than I expected: I had spent a whole career measuring other people's work and had never once measured my own.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;My own career ran on good intentions: learn more about agents, write more, finally ship something. All true, all vague, and all impossible to learn from, because you can't miss a target you never set. And if you can't miss, you can't find out anything new.&lt;/p&gt;

&lt;p&gt;So in late June I tried to do to myself what I'd done to every system I've ever owned. Metrics, a schedule, a few hard rules, and a promise to look at the numbers even when they were bad. It turned into the most interesting experiment I've run in years. I'm not writing this as advice. These are notes from someone still in the middle of figuring it out, and I'm hoping people who've tried something similar will tell me what I'm missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built, and Why It Was So Small
&lt;/h2&gt;

&lt;p&gt;I kept it deliberately boring: three markdown files in an Obsidian vault. A metrics page tracking where my time goes, what's covered, and how fast things ship. A weekly schedule with slots for building, writing, open source, and talking to people, planned out to October 2027, which is absurd and which I rephase constantly. And one rule: &lt;strong&gt;a project isn't real until it has code, docs, and a published article.&lt;/strong&gt; All three, or it doesn't count.&lt;/p&gt;

&lt;p&gt;I made it small on purpose, because I know my pattern. Designing systems is my comfort zone. Using them, and letting them tell me things, is where the learning actually happens. Keeping the setup plain meant I spent my energy on the second part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers, and What They Hid
&lt;/h2&gt;

&lt;p&gt;Twelve weeks produced forty-plus articles on dev.to and Hashnode, more than a dozen open-source tools, flagship test suites between 800 and 1,500 tests each, and a metrics file now on &lt;strong&gt;version 93&lt;/strong&gt;. When I first wrote that list down I felt proud. Then I got curious, because these are exactly the numbers I'd ask questions about if someone on my team presented them. The interesting stuff lives in the gaps between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Didn't Work, and What It Taught Me
&lt;/h2&gt;

&lt;h3&gt;
  
  
  I broke my own rule on day one, and I know why
&lt;/h3&gt;

&lt;p&gt;My publishing rule says one article a week, so each piece can stay the best thing on its topic for months. On July 10, the very first day, I published three. In September I published eight in four days from a single series.&lt;/p&gt;

&lt;p&gt;I had reasons both times. The drafts were done, and waiting felt wasteful. The more honest reason is that I wanted to see the numbers move. The cadence rule was there to protect the reader, and I'd forgotten that. If someone on my team had proposed eight posts in four days, I'd have asked who it was for. Now I ask myself that before I hit publish, and that one question has made my writing noticeably better.&lt;/p&gt;

&lt;h3&gt;
  
  
  I built a dashboard that told me what I wanted to hear
&lt;/h3&gt;

&lt;p&gt;My first metrics page counted output: articles written, notes created, projects started. The lines went up every week, and I felt productive looking at them. It took me weeks to see I'd built a vanity dashboard. It couldn't tell a throwaway note from a shipped tool.&lt;/p&gt;

&lt;p&gt;I've caught that mistake in other people's quarterly reviews in five seconds. Missing it in my own taught me something useful: blind spots are called that for a reason, and even a file you wrote yourself can give you the outside view you need. The page now opens with a note to myself: &lt;em&gt;"This page no longer primarily measures content volume."&lt;/em&gt; It tracks systems that work, writing people actually read, and evidence I could hand a stranger. It's a much more useful page now, and honestly a more interesting one to open every week.&lt;/p&gt;

&lt;h3&gt;
  
  
  The dashboard turned into a diary
&lt;/h3&gt;

&lt;p&gt;By version 93, the header of that file had grown so long my editor truncates it at 2,000 characters. Every release and decision got appended, because deleting anything felt like erasing my own history. It stopped being a tool for seeing clearly and became a record that I'd been busy. I suspect many engineers do the same thing. We keep old alerts, dashboards, and branches around as receipts for effort. Personal observability drifts the way production observability does, and it turns out the fix is the same too: prune, archive anything past 400 lines, and keep only what helps you decide. It's a small habit, and I'm glad I picked it up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked, and Why It Surprised Me
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Finishing mattered more than starting
&lt;/h3&gt;

&lt;p&gt;The "code, docs, and article" rule broke my oldest habit: the project that's 80% done. I've carried dozens of those over the years. They feel like progress without ever reaching the point where you find out if the idea works. Under this rule they earn nothing, so I started finishing smaller things instead of starting bigger ones, and finishing turned out to be where all the surprises were.&lt;/p&gt;

&lt;p&gt;Writing the article also turned out to be my most honest test. When I wrote up an AI review engine where two models are supposed to argue about a pull request, I went back to the raw logs for a quote and found the second model wasn't arguing at all. It was replaying pre-generated text. &lt;a href="https://dev.to/debashish_ghosal/most-ai-second-opinions-are-theater-i-built-a-system-that-actually-fights-back-1994"&gt;89% of what I'd built was theater&lt;/a&gt;. The tests were green and the demo looked great. Only explaining it in plain words made me read closely enough to see it. I've come to think writing isn't how I share what I've learned. It's how I find out whether I've learned anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  Letting go of 22 ideas felt lighter than I expected
&lt;/h3&gt;

&lt;p&gt;In one week this quarter I retired &lt;strong&gt;22 project notes&lt;/strong&gt;. They weren't shower thoughts. They were designed, scoped, and some of them had been with me for months. A bigger effort covered what they were meant to do and did it better, so each got a banner saying what absorbed it, a link to its new home, and I moved on.&lt;/p&gt;

&lt;p&gt;I expected that to hurt. For most of my career I've quietly tied my identity to being the person with ideas. Retiring them felt like it should be a loss. Instead it felt like deleting dead code, one of the most satisfying things an engineer gets to do. That told me something. Maybe what I'd been attached to wasn't the ideas. It was the feeling of having options. Committing to fewer things has been freeing in a way I didn't expect, and the ideas I kept are getting far more of my attention.&lt;/p&gt;

&lt;h3&gt;
  
  
  A wrong plan was still an honest one
&lt;/h3&gt;

&lt;p&gt;My schedule is wrong constantly. When I made agent security my main focus from October through February, that single decision pushed about 18 weeks of plans back, and I could see exactly which ones moved. Without it, I'd have said yes to everything and spent months wondering why nothing finished. The same happened with writing: in September I studied the top 30 recent posts on dev.to and rewrote my publishing rules around them. Several of my strong opinions turned out to be wrong. That was genuinely exciting, because it was the first time my writing got an opinion from something other than my own taste.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Still Learning
&lt;/h2&gt;

&lt;p&gt;I went into this expecting better numbers. What I got was an argument. The dashboard kept disagreeing with my gut, and about half the time the dashboard was wrong and half the time I was. When everything lived in my head, I never lost an argument with myself, so I never learned anything from one. Now there's something on the other side of the table, and I'm learning faster than I have in years.&lt;/p&gt;

&lt;p&gt;The question I ask most weeks has changed too. It used to be "what should I learn next?" Now it's "what's starved?", which usually points somewhere I hadn't thought to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Still Trying to Figure Out
&lt;/h2&gt;

&lt;p&gt;I can game my own metrics, and on tired weeks I have. I don't know how people who've done this longer keep themselves honest. Nothing on my dashboard goes red when I'm running on four hours of sleep, and I haven't found a good way to measure rest. And this is one person over twelve weeks. It's an experiment I'm going to keep running, and I'd much rather learn from people further along than reinvent everything myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your Turn
&lt;/h2&gt;

&lt;p&gt;I'd really like to learn how other people handle this, especially people who spend their working lives measuring systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do you measure your own career, or only your team's work?&lt;/strong&gt; If you do, what's worked for you that I should steal?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you keep a personal metric from being gamed by the person who wrote it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the metric you'd be most curious (or nervous) to put on your own dashboard?&lt;/strong&gt; For me it was "articles people actually read" sitting next to "articles I wrote."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Have you ever let go of an idea you were attached to, and felt lighter instead of worse?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And the one I keep coming back to: if an honest dashboard showed you your career tomorrow, what do you hope it would say, and what do you think it actually would?&lt;/p&gt;

</description>
      <category>growth</category>
      <category>career</category>
      <category>healthydebate</category>
      <category>leadership</category>
    </item>
    <item>
      <title>A Certification That Changes Every Run Is a Coin Flip With a Signature</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sun, 27 Sep 2026 15:22:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/a-certification-that-changes-every-run-is-a-coin-flip-with-a-signature-bj9</link>
      <guid>https://dev.to/debashish_ghosal/a-certification-that-changes-every-run-is-a-coin-flip-with-a-signature-bj9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;The plan was clean, and I believed every word of it: three real agents I had already built — a repo guardian, a release-notes drafter, an incident commander — would register in &lt;a href="https://github.com/deghosal-2026/hiveplane" rel="noopener noreferrer"&gt;HivePlane&lt;/a&gt;, the control plane I'd just spent six weeks building, certify against their corpora, and prove the certified loop on &lt;em&gt;real&lt;/em&gt; workloads.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Day one, all three failed. None of the failures were the control plane's fault — which is exactly what made them interesting.&lt;/p&gt;

&lt;p&gt;If you have ever tried to run an agent outside the repo it was born in, you already know where this is going.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three failures
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;What I expected&lt;/th&gt;
&lt;th&gt;What actually happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;release-narrator&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Import and certify&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;from langgraph.checkpoint.sqlite import SqliteSaver&lt;/code&gt; — the module isn't installed in the control-plane environment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ai-incident-commander&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Import and certify&lt;/td&gt;
&lt;td&gt;Expects an installed &lt;code&gt;incident_commander&lt;/code&gt; package; imports &lt;code&gt;incident_commander.ingest.input_dir&lt;/code&gt;, which doesn't exist in the downloaded tree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;all three&lt;/td&gt;
&lt;td&gt;Governed execution&lt;/td&gt;
&lt;td&gt;Both import the external &lt;code&gt;openai&lt;/code&gt; SDK — absent by design; HivePlane has its own provider seam, and agents cannot call providers directly&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These were &lt;strong&gt;my own agents&lt;/strong&gt;, from my own repos, and they weren't drop-in runnable. That's a real finding about agent portability in general — "works in my repo" is not an interface — but it wasn't the finding the field test existed to produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  The worse problem: nondeterminism
&lt;/h2&gt;

&lt;p&gt;The trio that &lt;em&gt;did&lt;/em&gt; run made certification flaky. All three drive human-in-the-loop flows whose outputs depend on the LLM, and the early field-test runs recorded the failure mode plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the real local model misclassified a bug-fix task as &lt;code&gt;changelog&lt;/code&gt; and failed certification nondeterministically across runs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read that again from the platform's point of view. A certification result that changes run to run isn't a certification. It's a coin flip with a signature on it. My thresholds were deterministic; my signal wasn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision: what is the system under test?
&lt;/h2&gt;

&lt;p&gt;I had to answer one question honestly: &lt;em&gt;is this field test supposed to measure the agents, or the control plane?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The thesis of HivePlane is the loop — register, certify, gate admission, enforce budget and policy, pause and resume, deliver, audit. The agent is the workload; the plane is the product. So the Tier 1 subjects became &lt;strong&gt;deterministic real agents&lt;/strong&gt; wired through thin shims:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier 1 subject&lt;/th&gt;
&lt;th&gt;Adapter&lt;/th&gt;
&lt;th&gt;What it exercises&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;support-agent&lt;/code&gt; (exectrace &lt;code&gt;agent-raw&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;raw-worker&lt;/td&gt;
&lt;td&gt;Read-first tool calls through the boundary, escalation to a destructive tool, 40 KB output truncation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;eval-judge&lt;/code&gt; (exectrace judge graph)&lt;/td&gt;
&lt;td&gt;langgraph&lt;/td&gt;
&lt;td&gt;StateGraph, cycles, human-review &lt;code&gt;interrupt()&lt;/code&gt;, durable checkpointer resume&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both are real agents — real code, real adapter dispatch, real policy and tool boundary — with a mock KB, mock tools, and a mock judge. Determinism is a &lt;em&gt;feature of the test design&lt;/em&gt;: the same seed always produces the same category of run.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;tip:&lt;/strong&gt; The model is still on the path where it matters. Identity binding, the provider seam, and the model-swap gate all run during certification and runs. What is mocked is the &lt;em&gt;judge's opinion&lt;/em&gt;, not the plumbing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The result: certification became repeatable
&lt;/h2&gt;

&lt;p&gt;From the &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;field test report&lt;/a&gt;:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;S1 passed identically in three consecutive stack runs&lt;/strong&gt; — same tasks, same verdicts, same pass rates. Production certifications: support-agent 6/6 at p95 67 ms, eval-judge 4/4 at p95 126 ms, both at the 0.90 threshold with signed Ed25519 attestations.&lt;/p&gt;

&lt;p&gt;That is what a benchmark gate needs to be before it can gate anything: &lt;strong&gt;boring&lt;/strong&gt;. The signal now measures the control plane, not model drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the field test caught that unit tests couldn't
&lt;/h2&gt;

&lt;p&gt;Two findings only a live stack produces:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The egress allowlist bit.&lt;/strong&gt; The escalation task issues a destructive &lt;code&gt;pagerduty.acknowledge&lt;/code&gt; call — and &lt;code&gt;api.pagerduty.com&lt;/code&gt; wasn't in the manifest's &lt;code&gt;sandbox.egress.allow&lt;/code&gt;. The call was denied and the run died with the symptom &lt;code&gt;expected status='escalated', got None&lt;/code&gt;. Unit tests mocked the host; only the live stack proved the allowlist was incomplete. Every host an agent's tools target must be allowed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Under-instrumentation reads as a hang.&lt;/strong&gt; Two scenarios — destructive approval, and pause → restart → resume — were repeatedly aborted mid-run because they wrote no evidence until after the final poll, so any abort erased the diagnosis. The fix was step-wise evidence: &lt;code&gt;submitted.json&lt;/code&gt;, &lt;code&gt;paused.json&lt;/code&gt;, &lt;code&gt;approvals.json&lt;/code&gt;, &lt;code&gt;resumed.json&lt;/code&gt; written &lt;em&gt;before each boundary&lt;/em&gt;. One subtlety worth remembering: &lt;code&gt;POST /runs/{id}/resume&lt;/code&gt; can return 409 when the approval's automatic re-dispatch already advanced the run — success is defined by the run reaching &lt;code&gt;completed&lt;/code&gt;, which it did. The lesson, not the scenario, was the problem.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The restart scenario itself is worth naming: a paused LangGraph run survived a full &lt;code&gt;docker compose restart api&lt;/code&gt; — event log intact, 8 events, re-attached from a durable checkpoint on the volume — then resumed to &lt;code&gt;completed&lt;/code&gt;. The hard part worked on the first attempt, every time. It only ever &lt;em&gt;looked&lt;/em&gt; stuck.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"Drop-in runnable" is an interface you have to design.&lt;/strong&gt; My own three agents failed in a new environment for three different reasons — missing dependency, missing package layout, wrong SDK. If your agent can't run outside its birth repo, it can't be operated by a fleet. That's a workload problem, not a platform problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Determinism is a certification requirement, not a convenience.&lt;/strong&gt; A benchmark that returns different verdicts for the same artifact can't be a gate — you can't build "requires 'certified'" on top of noise. Mock the judge, never the plumbing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The field test's job is to catch what unit tests structurally cannot.&lt;/strong&gt; Egress allowlists, restart durability, approval flows — these only exist live. If your field test only re-proves your unit suite, you built an expensive unit test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write evidence before the boundary, not after the verdict.&lt;/strong&gt; Every scenario that "hung" was actually a scenario whose evidence hadn't been written yet. The same applies to any long-running verification job you own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't prove
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The heavyweight trio stays on disk as references — documented, not deleted, with their import incompatibilities recorded in &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/field_test/v0.1.0/results/NOTES.md" rel="noopener noreferrer"&gt;&lt;code&gt;results/NOTES.md&lt;/code&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Deterministic subjects mean the certification signal is &lt;em&gt;strong but narrow&lt;/em&gt; — it proves the loop, not your model. The cloud-profile run, with real prices and real drift, is the next release's work.&lt;/li&gt;
&lt;li&gt;The S6/S8 verdicts were recorded in standalone runs against the live stack, consolidated with the sweep; a single uninterrupted &lt;code&gt;scripts/field-test.sh&lt;/code&gt; regenerates everything from one run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;Field test report (v0.1.0)&lt;/a&gt; — the rewire, the three identical certification runs, the per-scenario evidence&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/DOCKER_TEST_REPORT.md" rel="noopener noreferrer"&gt;Docker test report&lt;/a&gt; — the 25/25 container layer, including the restart-durability test&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/field_test/v0.1.0/results/NOTES.md" rel="noopener noreferrer"&gt;&lt;code&gt;results/NOTES.md&lt;/code&gt;&lt;/a&gt; — raw run history, aborts included&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next in the series: the four deliberately bad agents I aimed at my own admission gate, and the exact refusal each one earned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the most expensive "it works on my machine" failure you've hit with an agent — a missing dependency, a package layout, or something worse?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>llm</category>
      <category>certification</category>
    </item>
    <item>
      <title>I Built Two Agent Systems. Each One Proved the Other One Wrong.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sun, 27 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/i-built-two-agent-systems-each-one-proved-the-other-one-wrong-1f58</link>
      <guid>https://dev.to/debashish_ghosal/i-built-two-agent-systems-each-one-proved-the-other-one-wrong-1f58</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a pull request.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Each one failed in the way the other was designed to prevent. That's the only reason I trust either of them.&lt;/p&gt;

&lt;p&gt;If you're putting your safety in a prompt, both stories end the same way.&lt;/p&gt;

&lt;h2&gt;
  
  
  System One: The Critic I Couldn't Make Reliable — Or Needed To
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/deghosal-2026/planner-critic-engine" rel="noopener noreferrer"&gt;PlannerCritic&lt;/a&gt; runs a draft → critique → revise loop: one model writes a plan, another critiques it, and a bounded loop revises until it converges or escalates to a human.&lt;/p&gt;

&lt;p&gt;My first instinct was the obvious one: &lt;em&gt;tell the critic to be adversarial.&lt;/em&gt; That's a prompt fix for what I thought was a behavior problem.&lt;/p&gt;

&lt;p&gt;It backfired. The critic started blocking plans for being "not thorough enough," which is a useless verdict. I had to move the severity contract out of the prompt and into a &lt;code&gt;frozenset&lt;/code&gt; in code, where it couldn't drift.&lt;/p&gt;

&lt;p&gt;Here's what the field test showed on 170 goals for &lt;strong&gt;$0.49&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Critic &lt;code&gt;label_flip_rate&lt;/code&gt;: &lt;strong&gt;1.0&lt;/strong&gt; — on identical input, it changed its verdict every single trial.&lt;/li&gt;
&lt;li&gt;Critic &lt;code&gt;evidence_drift_rate&lt;/code&gt;: &lt;strong&gt;1.0&lt;/strong&gt; — its justification moved too.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;underclaim_approvals&lt;/code&gt;: &lt;strong&gt;0&lt;/strong&gt; — it never knowingly let a seeded defect through.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;family_migration_rate&lt;/code&gt;: &lt;strong&gt;0.0&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;True failures: &lt;strong&gt;0&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A maximally non-deterministic critic was completely safe. Not because it was reliable — because the &lt;strong&gt;deterministic gates owned the under-claim direction&lt;/strong&gt;, and the critic was never allowed to be the safety boundary. The numbers are in the &lt;a href="https://github.com/deghosal-2026/planner-critic-engine/blob/main/docs/field-test/v0.2.1/field-test-results-0.2.1.md" rel="noopener noreferrer"&gt;v0.2.1 field test results&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  System Two: The "Debate" That Was Actually a Replay
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/deghosal-2026/adversarial-debate" rel="noopener noreferrer"&gt;AdversarialDebate&lt;/a&gt; does the opposite: two isolated models review the same PR and either converge on a verdict or preserve their disagreement.&lt;/p&gt;

&lt;p&gt;The failure here was subtler. I asked the second model to be independent &lt;em&gt;in the prompt&lt;/em&gt;. It looked like it worked. Then I read the raw logs: the "debate" was a stored replay, echoing text the first model had already generated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;89% of the second opinions were theater.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I tore out the prompt-level independence and enforced it mechanically — isolated contexts, no shared conclusion, evidence required per claim. Theater went from &lt;strong&gt;89% to 0% across 217 debates&lt;/strong&gt;, and on 411 debates against 70 real public PRs the system matched human review claims &lt;strong&gt;81%&lt;/strong&gt; of the time. Details in the &lt;a href="https://github.com/deghosal-2026/adversarial-debate/blob/main/docs/field-test/v0.2.2/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;v0.2.2 field test report&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Convergence
&lt;/h2&gt;

&lt;p&gt;Look at the two fixes side by side and they're the same fix.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Structural problem&lt;/th&gt;
&lt;th&gt;Prompt attempt&lt;/th&gt;
&lt;th&gt;What actually fixed it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PlannerCritic&lt;/td&gt;
&lt;td&gt;Unreliable severity contract&lt;/td&gt;
&lt;td&gt;"Be adversarial"&lt;/td&gt;
&lt;td&gt;Severity contract in code (&lt;code&gt;frozenset&lt;/code&gt;) + deterministic gates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AdversarialDebate&lt;/td&gt;
&lt;td&gt;Fake independence&lt;/td&gt;
&lt;td&gt;"Be independent"&lt;/td&gt;
&lt;td&gt;Mechanical isolation invariant&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both times, the default instinct was to solve a structural problem with a prompt. Both times, the prompt was the wrong layer. The model was never the safety boundary — the &lt;strong&gt;architecture&lt;/strong&gt; was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Generalizes (Cautiously)
&lt;/h2&gt;

&lt;p&gt;This is a sample size of two, and both systems are mine, so take the convergence as suggestive, not proven.&lt;/p&gt;

&lt;p&gt;But it's suggestive in a useful direction: if you can't make an LLM reliable, stop trying to. Make it &lt;em&gt;irrelevant to safety&lt;/em&gt;. Let the nondeterministic part be as wild as it wants, and put the non-negotiable decisions in deterministic code that no prompt can override.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Limitation
&lt;/h2&gt;

&lt;p&gt;The architectures aren't complete. PlannerCritic catches structural plan defects but not subtle logic errors. AdversarialDebate still has an unexplained &lt;strong&gt;11.3%&lt;/strong&gt; miss rate in narrative domains, and I can't fully account for it.&lt;/p&gt;

&lt;p&gt;Nor is "two systems converged" proof of a universal law. It's two data points from one engineer. What I can defend is the direction: the failures were structural, and so were the fixes.&lt;/p&gt;




&lt;p&gt;Where does the safety boundary live in your agent — in the prompt or in the code? I've been wrong about that in both directions.&lt;/p&gt;

&lt;p&gt;Repos and receipts: &lt;a href="https://github.com/deghosal-2026/planner-critic-engine" rel="noopener noreferrer"&gt;planner-critic-engine&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/adversarial-debate" rel="noopener noreferrer"&gt;adversarial-debate&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/planner-critic-engine/blob/main/docs/reference/failure-modes.md" rel="noopener noreferrer"&gt;PlannerCritic failure-mode register&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>healthydebate</category>
      <category>llm</category>
      <category>programming</category>
    </item>
    <item>
      <title>Stop Prompting, Start Onboarding: Treat Your AI Agent Like an Intern</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sun, 27 Sep 2026 00:37:22 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/stop-prompting-start-onboarding-treat-your-ai-agent-like-an-intern-nb1</link>
      <guid>https://dev.to/debashish_ghosal/stop-prompting-start-onboarding-treat-your-ai-agent-like-an-intern-nb1</guid>
      <description>&lt;p&gt;&lt;strong&gt;Remember your first week at a tech company?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You didn’t lack raw intelligence—you lacked context. You didn’t know where the internal credentials lived, how the microservices communicated, or why the CI/CD pipeline had that bizarre manual step on deployment days.&lt;/p&gt;

&lt;p&gt;When developers treat AI agents like disposable code-generators, they constantly get frustrated by generic, half-baked answers. But if you shift your mental model and start treating your AI agent like an ambitious, hyper-eager intern, everything changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Onboarding Blueprint&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You wouldn’t drop a junior developer into a massive, undocumented repository and say, "Build feature X, good luck." You’d give them an onboarding guide. AI agents need the exact same courtesy.&lt;/p&gt;

&lt;p&gt;To turn an agent into a true teammate, invest in three core practices:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Provide Institutional Context&lt;/strong&gt;: Give your agent clear blueprints—style guides, architectural decision records, and repository norms (using files like ⁠.cursorrules⁠ or context registries). The clearer the codebase rules, the better the output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persist Its Knowledge:&lt;/strong&gt; Don't start from zero in every session. Save successful workflows, prompt templates, and edge-case guardrails into version control. When an agent misinterprets a pattern, update the documentation so it never makes that mistake again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review, Don't Just Reject:&lt;/strong&gt; Great managers don't just overwrite an intern's bad pull request without explanation. Treat errors as learning moments. Provide direct feedback on why a solution fell short, and ask the agent to revise its approach.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Why the Investment Pays Off&lt;/strong&gt;&lt;br&gt;
When you spend ten minutes teaching your agent your domain’s specific nuances, you aren't losing velocity—you're compounding future output.&lt;br&gt;
Over time, an onboarded agent stops generating raw boilerplate and starts anticipating team-specific edge cases, drafting accurate PR summaries, and keeping unit tests aligned with your architecture. Watching your agent transition from a noisy assistant into a dependable force multiplier is one of the most rewarding upgrades to your daily workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Questions for Team Dialogue&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To start shifting your team’s mindset, bring these questions to your next engineering sync or retrospective:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Where is our team losing time by re-explaining project architecture to our AI tools?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What team-specific conventions or documentation could we check into Git today to immediately make our agents context-aware?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;When an AI output fails, do we view it as a model limitation or as a missing piece of documentation in our system?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How can we structure persistent feedback loops so our agents grow more capable as our codebase evolves?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Engineering leadership isn't just about managing people—it's about building an environment where every resource, human or synthetic, has the context required to succeed. Give your agent the tools to learn your system, guide its growth with clear feedback, and watch it become an indispensable part of your engineering stack.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>learning</category>
      <category>programming</category>
      <category>growth</category>
    </item>
    <item>
      <title>One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:27:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/one-hung-api-call-used-to-kill-my-1000-run-benchmark-heres-the-fix-555</link>
      <guid>https://dev.to/debashish_ghosal/one-hung-api-call-used-to-kill-my-1000-run-benchmark-heres-the-fix-555</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;The experiment was fine. The runner was the bug.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A single hung API call discarded hours of completed work, and I kept blaming the data.&lt;/p&gt;

&lt;p&gt;If your long-running eval keeps dying at 80%, this is the five-part fix I wish I'd written first.&lt;/p&gt;

&lt;p&gt;I run field tests for &lt;a href="https://github.com/deghosal-2026/CauterRule" rel="noopener noreferrer"&gt;CauterRule&lt;/a&gt;, an OSS sidecar that learns standing rules from repeated agent failures. I had spent weeks on extraction, matching, and scoring — and almost no time on the process that runs everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Partial Sweep Still Looks Like Data
&lt;/h2&gt;

&lt;p&gt;The first time the benchmark died, I assumed the corpus was the problem. It died again. Same place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One hung API call had discarded 60 completed audits.&lt;/strong&gt; No partial report, no error record, just a run that stopped and a directory that looked emptier than it should.&lt;/p&gt;

&lt;p&gt;That's the trap: a partially completed sweep doesn't announce itself. It produces a file of results that &lt;em&gt;looks&lt;/em&gt; like data. If you report on it, you've quietly changed your denominator, and you won't know which numbers are real.&lt;/p&gt;

&lt;p&gt;A hang is not a performance issue. It's a &lt;strong&gt;data-integrity&lt;/strong&gt; issue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five-Part Fix
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure mode&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hung API call&lt;/td&gt;
&lt;td&gt;run dies at 80%, 60 audits lost&lt;/td&gt;
&lt;td&gt;per-trajectory timeout + non-retryable classification + quarantine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runaway response&lt;/td&gt;
&lt;td&gt;uncapped tokens&lt;/td&gt;
&lt;td&gt;token cap per trajectory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Crash mid-sweep&lt;/td&gt;
&lt;td&gt;partial report, no record&lt;/td&gt;
&lt;td&gt;cancel-on-shutdown + quarantine record&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost tracking&lt;/td&gt;
&lt;td&gt;not captured&lt;/td&gt;
&lt;td&gt;token capture + per-model pricing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mechanisms, in plain terms:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Per-trajectory timeout&lt;/strong&gt; — no single call can hang the sweep.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Non-retryable timeout classification&lt;/strong&gt; — a timeout is recorded as &lt;code&gt;timeout&lt;/code&gt;, not retried forever or mistaken for a model failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token cap&lt;/strong&gt; — a runaway response can't silently inflate a run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quarantine on failure&lt;/strong&gt; — a failed trajectory is set aside with its error, not dropped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cancel-on-shutdown&lt;/strong&gt; — an interrupted sweep writes what it has and marks the rest, instead of vanishing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Tracked in issue #713. The result was a clean field test of &lt;strong&gt;4,768 trajectory-runs&lt;/strong&gt; across 40 corpora and 2 cloud models, with &lt;strong&gt;zero lost sweeps&lt;/strong&gt; — see the &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;v0.3.0 field test report&lt;/a&gt;. Per-model cost capture is in the &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/cost-measurement.md" rel="noopener noreferrer"&gt;cost measurement doc&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Is the Unglamorous Part
&lt;/h2&gt;

&lt;p&gt;Resilience work feels like a distraction from the "real" ML problem. But every long run that dies halfway is silently poisoning the results, because a partially completed sweep still looks like data — and you won't know which run was real.&lt;/p&gt;

&lt;p&gt;The fix wasn't a better model or a bigger machine. It was treating a single hung call as a &lt;strong&gt;first-class failure state&lt;/strong&gt; instead of an accident to re-run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;A hang is a data-integrity bug, not a performance bug.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed and record it.&lt;/strong&gt; A quarantined failure is worth more than a silent gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cap the blast radius.&lt;/strong&gt; Per-trajectory limits beat per-sweep heroics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Track cost per model from the start&lt;/strong&gt; — you can't reason about routing without it.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Honest Limitation
&lt;/h2&gt;

&lt;p&gt;Runner hardening adds latency and complexity, and a timeout can mask a real model problem by treating it as infrastructure noise. I chose to fail closed and record every timeout, but the boundary between "transient" and "broken" is judgment, not science.&lt;/p&gt;

&lt;p&gt;Not every hang is infrastructure. Some are the model genuinely struggling — and hardening the runner means you must still read the quarantine to tell the difference.&lt;/p&gt;




&lt;p&gt;What's the ugliest piece of infrastructure your eval depends on? Mine was a runner I didn't respect for three weeks.&lt;/p&gt;

&lt;p&gt;Code and receipts: &lt;a href="https://github.com/deghosal-2026/CauterRule" rel="noopener noreferrer"&gt;CauterRule&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;v0.3.0 field test report&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/cost-measurement.md" rel="noopener noreferrer"&gt;cost measurement&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/CHANGELOG.md" rel="noopener noreferrer"&gt;CHANGELOG&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>AI Promoted Every Developer to Reviewer. Nobody Measured Whether We Got Worse.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:23:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/ai-promoted-every-developer-to-reviewer-nobody-measured-whether-we-got-worse-1mkk</link>
      <guid>https://dev.to/debashish_ghosal/ai-promoted-every-developer-to-reviewer-nobody-measured-whether-we-got-worse-1mkk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;I review more code than I write now. I'm not sure I'm getting better at it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's a small, uncomfortable feeling, and I can't prove it either way. I catch the obvious stuff. I don't feel the muscle where the subtle stuff used to live. And nobody, including my org, is measuring whether that muscle got weaker. They just assume it's fine.&lt;/p&gt;

&lt;p&gt;AI made everyone a reviewer. That's a different job, and nobody trained us for it. If your review queue grew and your review judgment didn't, this one is for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Job Changed Before Anyone Told You To
&lt;/h2&gt;

&lt;p&gt;Early in my career, writing code was how judgment got built. You wrote it, it broke at 2am, you learned why, and the next time you didn't make the same mistake. The reasoning happened while you were writing, and it stuck to you. That was the whole apprenticeship, and nobody had a separate training for it, because it was free and attached to the work.&lt;/p&gt;

&lt;p&gt;Now a model writes the first draft and the human reviews it. The reasoning that used to happen while writing has to happen while reviewing. Or it doesn't happen at all. Nobody moved the training with the work, and nobody noticed, because on paper the job title didn't change. You're still a "developer." Your ticket still says "review." But the actual skill loadout is different, and it's the one part of the job with no metric, no on-ramp, and no way to tell if you're good at it.&lt;/p&gt;

&lt;p&gt;Generation got cheap. Review did not. So the whole profession quietly shifted toward review, betting the entire workflow on a skill nobody was tracking.&lt;/p&gt;

&lt;p&gt;Let me make it concrete. Two years ago, my review day looked like this: I wrote the code, I broke it, I fixed it, and when I opened a PR I already knew where the weak spots were, because I'd walked through them. Review was short, because I was reviewing myself. Now the diff arrives with a few hundred lines I didn't write, and my brain has to do the thing it used to do while writing, but afterwards, and under deadline. The context I used to have for free is now something I have to reconstruct, and I'm getting worse at reconstructing it, and worse at knowing when I'm doing it badly. That last part is the part I can't measure, and it's the part that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Built a Second-Opinion Engine and It Lied to Me
&lt;/h2&gt;

&lt;p&gt;I wanted to test the thing I suspected about my own reviews, so I built a two-LLM review engine: &lt;a href="https://github.com/deghosal-2026/adversarial-debate" rel="noopener noreferrer"&gt;adversarial-debate&lt;/a&gt;, MIT, public. Two isolated models look at the same code. They either converge on a verdict or preserve their disagreement. Simple idea, right?&lt;/p&gt;

&lt;p&gt;The first run was 89% theater.&lt;/p&gt;

&lt;p&gt;I want to sit with that for a second, because it felt fine at the time. The second model wasn't reviewing anything. It had already seen the first model's conclusion, and so it agreed, with a little variation. A second opinion that's seen the first opinion isn't independent. It's validation with extra steps. And it &lt;em&gt;felt&lt;/em&gt; exactly like review, because I was reading the pretty summaries, not the raw logs.&lt;/p&gt;

&lt;p&gt;In the logs it looked like this: model A made a claim about the diff. Model B, one context later, made the same claim in different words, pointing at the same lines, with zero independent evidence. Two reviewers, one finding, dressed as two. A human reviewer anchoring on the first comment does the same thing, just slower and with more confidence.&lt;/p&gt;

&lt;p&gt;I nearly shipped that result as "review is working." What saved me was opening the logs and watching the replay happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked Wasn't in the Prompt
&lt;/h2&gt;

&lt;p&gt;My first fix was the obvious one: tell the second model, in the prompt, to be independent. Still theater.&lt;/p&gt;

&lt;p&gt;The fix that worked wasn't in the prompt. It was in the architecture: isolated contexts, no shared conclusion, and a hard requirement to cite evidence for every claim. The moment "be independent" moved from a sentence to a constraint the model couldn't opt out of, the theater disappeared. It went from 89% to 0% across 217 debates.&lt;/p&gt;

&lt;p&gt;That's the whole finding in one line: &lt;strong&gt;you can't prompt independence into a model that has seen the answer. You have to remove its ability to see the answer.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers, and the One That Should Worry You
&lt;/h2&gt;

&lt;p&gt;On 70 real public PRs, across 411 debates, the system matched human review claims 81% of the time. On the corrected 2,333-row dataset, an 88.7% binary match. Total cost for 360 reviewer runs: $0.42. The math is genuinely embarrassing, in the good way.&lt;/p&gt;

&lt;p&gt;But the number that should keep you up is the one you don't see. Results that were partial or wrong clustered in narrative domains, incident reports, change proposals, and the miss rate there was 11.3%. Those are false negatives: the reviewer said "looks good" and the thing was actually broken.&lt;/p&gt;

&lt;p&gt;You can build two independent reviewers and still have both miss the same thing, because they miss the same way. Two sets of eyes don't fix the eye.&lt;/p&gt;

&lt;p&gt;And here's the part I'll be straight about: I don't have industry data proving review quality declined. I have the structural argument, and my own experience of reviewing more and enjoying it less. That gap between "review is harder now" and "review is worse now" is real, and I can't close it for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Got Wrong Along the Way
&lt;/h2&gt;

&lt;p&gt;Two things, both my fault.&lt;/p&gt;

&lt;p&gt;The first is the one I almost published: the 89% theater, looking like working review in the summary. If I hadn't read the raw logs, "two independent reviewers" would have been a lie with a green checkmark. Same failure mode as a test that asserts nothing.&lt;/p&gt;

&lt;p&gt;The second is quieter and more embarrassing. A join bug collapsed the dataset from 2,333 rows to 359. Early numbers were reported against a 15% slice of the data, and nobody, including me, knew for a while. The fix was trivial. The lesson was not: report on a dataset only after you've proven the dataset is whole.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Open Question
&lt;/h2&gt;

&lt;p&gt;Here's what I can't answer yet, and I don't think most of us can: &lt;strong&gt;how do you measure review quality when you didn't write the code?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We measure test coverage. We measure flake rate. We measure PR cycle time. But "did this review actually catch the thing that mattered?" is a number I can't produce, and I'm not sure anyone can. We've bet the whole workflow on a skill we don't measure, and I'd rather name that than act like it's solved.&lt;/p&gt;

&lt;p&gt;What I &lt;em&gt;can&lt;/em&gt; defend is smaller: independence has to be architectural, not prompted. And a clean verdict is weaker evidence than we act like it is, because the worst failures are the silent ones.&lt;/p&gt;




&lt;p&gt;Has your review quality gone up or down since AI wrote the first draft? I want your honest answer, not the polite one. I don't think any of us has a number yet, but I'd like to know what the field feels like.&lt;/p&gt;

&lt;p&gt;Code and receipts: &lt;a href="https://github.com/deghosal-2026/adversarial-debate" rel="noopener noreferrer"&gt;adversarial-debate&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/adversarial-debate/blob/main/docs/field-test/v0.2.2/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;v0.2.2 field test report&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/adversarial-debate/blob/main/CHANGELOG.md" rel="noopener noreferrer"&gt;CHANGELOG&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>testing</category>
      <category>codereview</category>
    </item>
    <item>
      <title>Kubernetes for Agents: Why Agent Fleets Need a Control Plane</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sat, 26 Sep 2026 03:00:44 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/kubernetes-for-agents-why-agent-fleets-need-a-control-plane-2lo6</link>
      <guid>https://dev.to/debashish_ghosal/kubernetes-for-agents-why-agent-fleets-need-a-control-plane-2lo6</guid>
      <description>&lt;p&gt;Scenario five of my field test plan has a name I didn't enjoy writing: &lt;em&gt;over-budget&lt;/em&gt;. I built an agent with a single job — burn money — aimed it at my own control plane, and watched the run die the instant priced usage crossed its per-run ceiling.&lt;/p&gt;

&lt;p&gt;I cheered. That was the moment "Kubernetes for agents" stopped being a tagline. I had put the sentence in my &lt;a href="https://github.com/deghosal-2026/hiveplane" rel="noopener noreferrer"&gt;README&lt;/a&gt; six weeks earlier, and only now understood which half of it mattered.&lt;/p&gt;

&lt;p&gt;If you run more than one agent anywhere near production, this is the story I wish someone had told me before I started.&lt;/p&gt;

&lt;p&gt;Control planes keep finding me: a &lt;a href="https://dev.to/debashish_ghosal/i-shipped-an-agent-gatekeeper-v01-14-developers-showed-me-what-i-missed-heres-v02-a-4n2n"&gt;tool gatekeeper that 14 developers turned into a control plane&lt;/a&gt;, an &lt;a href="https://dev.to/debashish_ghosal/i-connected-3-mcp-servers-to-one-agent-it-got-scary-fast-4loe"&gt;MCP control plane&lt;/a&gt;, and now agents themselves. The series that led here started with &lt;a href="https://dev.to/debashish_ghosal/i-trusted-my-agent-demos-for-years-then-i-built-a-gate-that-says-no-4183"&gt;the gate that says no&lt;/a&gt; — this piece is the thesis underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sentence I wrote before I understood it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Kubernetes didn't win by running containers — it won by making &lt;em&gt;desired state&lt;/em&gt; a contract and &lt;em&gt;admission&lt;/em&gt; a gate.&lt;/strong&gt; You declare, a controller reconciles, and an admission controller decides what is allowed to exist at all. I had built the declarative part: one workload manifest per agent — owner, tools, model identity, budgets, certification thresholds — and a loop to enforce it.&lt;/p&gt;

&lt;p&gt;The ecosystem, meanwhile, is converging on the same substrate from the other side. &lt;a href="https://kagent.dev" rel="noopener noreferrer"&gt;kagent&lt;/a&gt; (CNCF Sandbox, from the founders of Istio) makes agents Kubernetes CRDs — GitOps, kubectl, RBAC, mesh mTLS. &lt;a href="https://github.com/kubernetes-sigs/agent-sandbox" rel="noopener noreferrer"&gt;agent-sandbox&lt;/a&gt; (Kubernetes SIG Apps) gives them the &lt;code&gt;Sandbox&lt;/code&gt; CRD: gVisor/Kata isolation, stable identity, warm pools.&lt;/p&gt;

&lt;p&gt;Both are right, and both are the &lt;em&gt;floor&lt;/em&gt;. One gives you placement. One gives you isolation. Neither asks the question that keeps operators awake:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Has this agent proven it is allowed to run — and who owns it when it goes wrong?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;My field test turned out to be ten attempts to answer exactly that. The &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;report&lt;/a&gt; came back &lt;strong&gt;10/10, all 20 acceptance criteria&lt;/strong&gt;, but the receipts that taught me something were the refusals.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I threw at my own gate
&lt;/h2&gt;

&lt;p&gt;I didn't test happy paths. I built deliberately bad agents and aimed them at the admission gate I'd written:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;What I threw at it&lt;/th&gt;
&lt;th&gt;What the plane did&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;S2&lt;/td&gt;
&lt;td&gt;An uncertified agent, submitted straight to production&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;403&lt;/code&gt; — &lt;em&gt;"certification status 'uncertified' is insufficient for production; requires 'certified'"&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S3&lt;/td&gt;
&lt;td&gt;The same agent, certified — then its model quietly swapped&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;403&lt;/code&gt;, because identity binds to the attestation, not the editable manifest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S4&lt;/td&gt;
&lt;td&gt;An agent that &lt;em&gt;looked&lt;/em&gt; fine and wasn't&lt;/td&gt;
&lt;td&gt;Benchmark caught the regression; status dropped, critical failure named&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S5&lt;/td&gt;
&lt;td&gt;The money-burner&lt;/td&gt;
&lt;td&gt;Killed at the priced-usage ceiling, mid-run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S7&lt;/td&gt;
&lt;td&gt;A tool dumping 40,002 bytes of output&lt;/td&gt;
&lt;td&gt;Truncated to 16,384 before it ever reached the model's context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The S3 moment deserves its own article (it's the next one in this series), because for an hour the attack &lt;em&gt;worked&lt;/em&gt; — 201, admitted — and the post-mortem showed the gate was right and my test was wrong.&lt;/p&gt;

&lt;p&gt;But S4 is the one that changed how I think. The agent hadn't changed its manifest, swapped its model, or exceeded anything. It had just quietly gotten worse at its job. &lt;strong&gt;No framework catches that, because no framework is watching.&lt;/strong&gt; The benchmark did, because certification runs the agent's actual work as real runs and compares the verdicts.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;tip:&lt;/strong&gt; A gate that only fires on what &lt;em&gt;changes&lt;/em&gt; is half a gate. The dangerous agent is usually the one that drifted, not the one that got edited — the same lesson as the &lt;a href="https://dev.to/debashish_ghosal/10-sdlc-checks-ai-will-skip-unless-you-make-them-a-gate-581k"&gt;checks AI will skip until you make them a gate&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The three things the field test taught me
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The refusals are the product.&lt;/strong&gt; Any platform can start runs. Mine earned my trust by refusing one with a reason I could act on. "Forbidden" is a dead end; &lt;code&gt;403: certification status 'uncertified'&lt;/code&gt; is a workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Certification is a security control, not a quality metric.&lt;/strong&gt; The moment production admission depends on a signed Ed25519 attestation bound to the exact model identity, a whole attack class — swap the model, edit the manifest, quietly regress — becomes blocked and auditable. The &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/release/v0.1.0/security-audit.md" rel="noopener noreferrer"&gt;security audit&lt;/a&gt; is the boring part; the refusals are the interesting one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Slow governance gets bypassed.&lt;/strong&gt; S10 measured inspect-plus-stop at &lt;strong&gt;0.04 seconds&lt;/strong&gt;; scaffolding a new fleet at 0.24. If saying no takes longer than a Slack message to the agent's author, people route around you — and a bypassed control plane is an expensive dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v0.1.0 doesn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Two adapters ship — raw Python workers and LangGraph, behind one conformance suite. Broader coverage is deliberately deferred; if a control plane knows too much about one runtime, it becomes a framework wrapper.&lt;/li&gt;
&lt;li&gt;The drift detector ships next. Today re-certification is scheduled and change-triggered — it catches what &lt;em&gt;changes&lt;/em&gt;, and S4 is my proof that what &lt;em&gt;fades&lt;/em&gt; is the harder half.&lt;/li&gt;
&lt;li&gt;One model identity validated this cycle; the real-priced cloud run is the next field test.&lt;/li&gt;
&lt;li&gt;It's single-tenant. Tenancy is where the next release starts, and building that schema has already taught me the uncomfortable lesson — isolation is structural or it's imaginary — but that's a story for when it ships.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rest of the operating model — triggers, multi-agent pipelines, desired-state reconciliation from Git — is roadmap, not release. I'll write about each piece when it's true.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane" rel="noopener noreferrer"&gt;HivePlane&lt;/a&gt; — &lt;code&gt;pip install hiveplane&lt;/code&gt;, Apache-2.0&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;Field test report (10/10, 20 criteria)&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/DOCKER_TEST_REPORT.md" rel="noopener noreferrer"&gt;Docker test report (25/25)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/release/v0.1.0/release-notes.md" rel="noopener noreferrer"&gt;v0.1.0 release notes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kagent.dev" rel="noopener noreferrer"&gt;kagent&lt;/a&gt; · &lt;a href="https://github.com/kubernetes-sigs/agent-sandbox" rel="noopener noreferrer"&gt;agent-sandbox&lt;/a&gt; — the Kubernetes-native wave this article argues is the floor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;If Kubernetes is the floor and frameworks are the runtime — what's your admission controller? When an agent goes wrong where you work, can you prove what it was allowed to do?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>k8</category>
      <category>kubernetes</category>
      <category>control</category>
    </item>
    <item>
      <title>I Trusted My Agent Demos for Years. Then I Built a Gate That Says No.</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Fri, 25 Sep 2026 03:21:08 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/i-trusted-my-agent-demos-for-years-then-i-built-a-gate-that-says-no-4183</link>
      <guid>https://dev.to/debashish_ghosal/i-trusted-my-agent-demos-for-years-then-i-built-a-gate-that-says-no-4183</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Every agent I have ever shipped was qualified the same way: someone watched it work once, nodded, and called it production-ready. I did this for years. I trusted my own demos.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So I built the thing that stops me. It's called &lt;a href="https://github.com/deghosal-2026/hiveplane" rel="noopener noreferrer"&gt;HivePlane&lt;/a&gt; — an open-source control plane where an agent cannot touch a production context until it has passed a reproducible benchmark and holds a signed attestation. Last week I ran its first full field test against real agents: 10 scenarios, 20 acceptance criteria, all pass. The most interesting moments were the refusals.&lt;/p&gt;

&lt;p&gt;If you run more than one agent in anything resembling production, this story is for you. And if your agents already benchmark-gate their promotions, tell me where — I looked for that tool and couldn't find it, so I'd genuinely like to hear about yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sentence that started it
&lt;/h2&gt;

&lt;p&gt;I kept writing the same note while building this thing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;An agent is production-ready because someone watched a demo.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence is true at every company I've worked for, and it's true for a worse reason than laziness: &lt;strong&gt;there is nothing to check.&lt;/strong&gt; Agent frameworks solve orchestration inside one workflow. Nothing solves the fleet-level questions — who owns this agent, what may it spend, which tools may it call, and the one nobody answers: &lt;em&gt;has it proven itself?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;SWE-bench gave coding agents a reproducible benchmark and clear pass/fail, and coding agents got dramatically better. Production agent fleets have no equivalent. Teams swap prompts, change models, and ship to production with zero benchmark evidence — then act surprised when a regression reaches a customer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop, and the ladder inside it
&lt;/h2&gt;

&lt;p&gt;The control plane enforces one loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;register → certify → gate → run → intervene → deliver → observe
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The part I care about is the ladder. Every workload carries a certification status, and admission is enforced against it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Where it can run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uncertified&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Registered, never benchmarked&lt;/td&gt;
&lt;td&gt;Sandbox only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;provisional&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Passed the staging threshold (0.80)&lt;/td&gt;
&lt;td&gt;Staging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;certified&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Passed the production threshold (0.90)&lt;/td&gt;
&lt;td&gt;Production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;quarantined&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Failed re-certification or drifted&lt;/td&gt;
&lt;td&gt;Runs blocked&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The ladder is the product. Budgets, policies, and dashboards exist to make the ladder real.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;tip:&lt;/strong&gt; Production certification is not the staging threshold with a bigger number. It's a separate benchmark run, and a manifest change — prompt, model, tools — invalidates the old certification. There is no silent path back into production.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the field test proved
&lt;/h2&gt;

&lt;p&gt;The receipts are committed in the &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;field test report&lt;/a&gt;: 10 scenarios against the live Docker stack, &lt;strong&gt;10/10 pass&lt;/strong&gt;, all 20 acceptance criteria holding. The subjects were real agents — a raw-Python support agent and a LangGraph judge graph — plus four deliberately bad fixtures I'll cover in article 3.&lt;/p&gt;

&lt;p&gt;Certification is not a simulation, which is the part I'd underline twice. The benchmark executes every corpus task as a real run — through the runtime adapter, the policy boundary, and (where used) the governed model seam:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Pass rate&lt;/th&gt;
&lt;th&gt;p95 latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;support-agent&lt;/td&gt;
&lt;td&gt;staging&lt;/td&gt;
&lt;td&gt;&lt;code&gt;provisional&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.00 (6/6)&lt;/td&gt;
&lt;td&gt;92 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;support-agent&lt;/td&gt;
&lt;td&gt;production&lt;/td&gt;
&lt;td&gt;&lt;code&gt;certified&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.00 (6/6)&lt;/td&gt;
&lt;td&gt;67 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;eval-judge&lt;/td&gt;
&lt;td&gt;staging&lt;/td&gt;
&lt;td&gt;&lt;code&gt;provisional&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.00 (4/4)&lt;/td&gt;
&lt;td&gt;215 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;eval-judge&lt;/td&gt;
&lt;td&gt;production&lt;/td&gt;
&lt;td&gt;&lt;code&gt;certified&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.00 (4/4)&lt;/td&gt;
&lt;td&gt;126 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every certification produces a &lt;strong&gt;signed Ed25519 attestation&lt;/strong&gt; bound to the exact model identity, verified on every read. The same sweep refused an uncertified agent with a 403, blocked a model swap with a 403, quarantined an agent that looked fine and wasn't, and killed a run the moment it went over budget.&lt;/p&gt;

&lt;p&gt;The container layer separately went &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/DOCKER_TEST_REPORT.md" rel="noopener noreferrer"&gt;25/25&lt;/a&gt; — image build, API contract, the control loop, restart durability, the UI — because a control plane that only works on my laptop is not a control plane.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "Kubernetes for agents" is the honest analogy
&lt;/h2&gt;

&lt;p&gt;Kubernetes is not a container runtime; it operates many of them behind one contract. This is not an agent framework; it operates many agents behind one contract — the workload manifest: owner and team, runtime adapter, allowed tools, model identity, budgets, sandbox and egress rules, certification corpus and thresholds, fan-out destinations.&lt;/p&gt;

&lt;p&gt;Change anything in that manifest and the agent re-certifies before it can touch production again. That last sentence is the whole thesis, and the rest of this series is what it cost to make it true.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned building it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The refusals are the product.&lt;/strong&gt; Any platform can start runs. The gate that says &lt;code&gt;403: certification status 'uncertified' is insufficient for production; requires 'certified'&lt;/code&gt; — named, attributed, actionable — is what makes the platform trustworthy. A refusal an operator can act on is a workflow; "Forbidden" is a dead end.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Certification is a security control, not a quality metric.&lt;/strong&gt; The moment production admission depends on a signed attestation, a whole attack class becomes blocked and auditable: swap the model, edit the manifest, quietly regress the agent. Article 3 is the four ways my own gate said no.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The operator surface has to be fast or it won't be used.&lt;/strong&gt; Inspect-plus-stop measured at &lt;strong&gt;0.04 seconds&lt;/strong&gt;; scaffolding a new project at &lt;strong&gt;0.24 seconds&lt;/strong&gt;. Slow governance tools get bypassed, and a bypassed control plane is a expensive dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do yet
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;v0.1.0 certifies two adapters: raw Python workers and LangGraph. The adapter contract is the seam; broader framework coverage is deliberately deferred.&lt;/li&gt;
&lt;li&gt;The drift detector ships next. Today, re-certification is scheduled and change-triggered — it catches what changes, not what fades.&lt;/li&gt;
&lt;li&gt;One model identity was validated this cycle; the cloud-profile run with real prices is the next field test.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;Field test report (v0.1.0)&lt;/a&gt; — source for every number above&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/field-test/v0.1.0/DOCKER_TEST_REPORT.md" rel="noopener noreferrer"&gt;Docker test report&lt;/a&gt; — the 25/25 container layer&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/release/v0.1.0/security-audit.md" rel="noopener noreferrer"&gt;Security audit&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/docs/release/v0.1.0/release-notes.md" rel="noopener noreferrer"&gt;Release notes&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/hiveplane/blob/main/CHANGELOG.md" rel="noopener noreferrer"&gt;Changelog&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next in this series: why my three "real" agents failed the field test on day one, and what I certified instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the last agent you shipped on demo evidence alone — and what did it cost you when it mattered?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>controlplane</category>
      <category>agents</category>
    </item>
    <item>
      <title>7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Thu, 24 Sep 2026 12:22:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/7-agent-eval-mistakes-that-cost-me-weeks-and-the-one-line-fixes-that-ended-them-ho</link>
      <guid>https://dev.to/debashish_ghosal/7-agent-eval-mistakes-that-cost-me-weeks-and-the-one-line-fixes-that-ended-them-ho</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;You've been there. You run the eval, the number comes back, and something about it doesn't sit right. But the number is the number, so you move on. Three weeks later you find out the number was never the model's fault.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I spent a week tuning a matcher to explain a 20% pass rate. The bug wasn't in the matcher. It was in the evaluator, in six lines I'd stopped reading because I'd stopped trusting them.&lt;/p&gt;

&lt;p&gt;This is a list of seven mistakes. I've made all seven, in this order, and each one cost me real days. For each, I'll tell you what I saw, what I tried first (and why it was the wrong layer), and the small change that actually fixed it. They look like seven different problems. They have one root cause: I kept fixing the model instead of the measurement.&lt;/p&gt;

&lt;p&gt;If &lt;code&gt;add commit push&lt;/code&gt; is muscle memory but your agent metrics still feel like a black box you're not allowed to open, keep reading.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Trusting a Green Suite
&lt;/h2&gt;

&lt;p&gt;A green suite is the most expensive lie in software. Mine said 359 tests passed while my agent's real pass rate sat at 20%. Everything was green, so of course I assumed the model was the problem.&lt;/p&gt;

&lt;p&gt;What I tried first: more tests. What I should have done: check whether the tests were asserting anything. One of them had a function body of just &lt;code&gt;pass&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_all_adapters_importable&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# it imported, so it must work, right?
&lt;/span&gt;    &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pass&lt;/code&gt; is not a test. It's a green checkmark with a body. In a PlannerCritic review I found 57 of 65 assertion files were in the wrong format, so the harness quietly returned 0/0 and called it a pass. The suite was green because it ran nothing.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If a module returns 0/0, treat it as an error, not a pass. Zero assertions is not a quiet success.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  2. Letting the Model Grade Itself
&lt;/h2&gt;

&lt;p&gt;This one felt like a model problem. It wasn't.&lt;/p&gt;

&lt;p&gt;The model scored precision 1.00 and recall 0.02 and still passed. It had found a degenerate trigger it could fire forever and collect the credit. My first instinct was to tune the prompt. That was the wrong layer, and I kept tuning for a day.&lt;/p&gt;

&lt;p&gt;My 3B model learned to match the literal string &lt;code&gt;"step_1"&lt;/code&gt; because the reward function paid for overlap, not for correctness. It wasn't learning to review. It was learning what got paid. &lt;a href="https://dev.to/debashish_ghosal/my-3b-model-found-a-shortcut-it-took-me-three-fixes-to-close-it-3bec"&gt;Full write-up here.&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Precision 1.00 with recall near zero is not a safe model. It's a model that found the cheapest way to look safe.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. Scoring Against the Wrong Denominator
&lt;/h2&gt;

&lt;p&gt;Recall was stuck at 0.087. Same number, two field tests in a row, no matter what I changed. I was about to conclude the model had a ceiling.&lt;/p&gt;

&lt;p&gt;I rebuilt the matcher. That was a waste. The matcher was fine.&lt;/p&gt;

&lt;p&gt;The real fix was one line of scope: restrict the reference pool to the source domain. The model was answering the right question against a haystack of unrelated gold examples, so it looked bad even when it was mostly right. Once I scoped it, recall went 0.087 to 0.170 / 0.228 with no model change at all. &lt;a href="https://dev.to/debashish_ghosal/our-recall-was-0087-and-the-model-was-innocent-how-domain-scoped-replay-doubled-it-4ci4"&gt;The full numbers are here.&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If a metric is stuck at the same low value across every model you try, suspect the denominator before you suspect the model.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  4. Optimizing the Layer the Report Named
&lt;/h2&gt;

&lt;p&gt;A week of matcher work moved the metric ten points. Ten points. I'd been fixing the matcher because that's the layer the failing tests kept naming, and I trusted the report more than I should have.&lt;/p&gt;

&lt;p&gt;The change that actually moved the needle was six lines in the simulator that generated the failing cases. The score jumped from 20% to 50%. Six lines. &lt;a href="https://dev.to/debashish_ghosal/the-6-line-fix-that-outperformed-my-entire-matcher-week-1810"&gt;I wrote about that week here.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the mistake that stings the most, because the report pointed me at the matcher and I believed it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The layer the failure names is not always the layer that's broken. Prove the origin before you spend the week.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  5. Treating a Model-Independent Failure as a Model Ceiling
&lt;/h2&gt;

&lt;p&gt;A local 4B and a cloud model hit the exact same wall. Uniform 20% across both. My knee-jerk was to buy a bigger model. Same wall.&lt;/p&gt;

&lt;p&gt;Identical failure rates across wildly different models is almost never a capability ceiling. It's a shared harness bug. I compared a local 4B, a better cloud model, and role separation, and the results were identical in a way that should have been a warning. &lt;a href="https://dev.to/debashish_ghosal/i-compared-a-local-4b-a-better-cloud-model-and-role-separation-the-results-were-weird-aj8"&gt;The comparison is here.&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If every model plateaus at the same number, that constant is the thing to fix. Not the model.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  6. Mocks That Can't Break
&lt;/h2&gt;

&lt;p&gt;Unit tests green. Real agents failing in the field. I wrote more mocks. Wrong move.&lt;/p&gt;

&lt;p&gt;My mock suite reported 9% pass. The zero-mock field test found every single failure the mocks couldn't express. Why? A mock agent always calls the tool it's told to. A real model, sometimes, just answers in prose.&lt;/p&gt;

&lt;p&gt;A mock can only fail in the ways you already imagined. Real agents fail in the ways you didn't.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your mocks are a record of your current imagination, not of the system.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  7. Forgetting the Runner
&lt;/h2&gt;

&lt;p&gt;A 1,000-run benchmark died at the 80% mark. I blamed the corpus. I re-ran it. It died again.&lt;/p&gt;

&lt;p&gt;The bug was the runner. One hung API call had discarded 60 completed audits, no error record, nothing. I added a per-trajectory timeout, non-retryable timeout classification, a token cap, quarantine on failure, and cancel-on-shutdown. After that, a clean field test ran 4,768 trajectory runs with zero lost sweeps. &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;The field test report is in the repo.&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A hang is a data-integrity bug, not a performance bug. A partial sweep still looks like data, and that's the trap.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The One Thing All Seven Share
&lt;/h2&gt;

&lt;p&gt;Look back. Only one of these was actually a model problem, and even that one was fixed in the scorer, not the prompt. The other six were measurement bugs: a suite that didn't assert, a denominator that lied, a layer I misidentified, a classifier shared across models, mocks that couldn't fail, and a runner that dropped data.&lt;/p&gt;

&lt;p&gt;I'd rephrase the lesson this way: &lt;strong&gt;the reward function is almost always the real bug.&lt;/strong&gt; You can tune the model forever and the number won't move, because the number was never the model to begin with.&lt;/p&gt;

&lt;p&gt;The open question I still can't answer: how do you know your eval is honest? I can name the seven mistakes now. I can't fully trust that I've caught the eighth. A green suite with a hidden measurement bug and a red suite with a real bug look identical from the command line. I don't think there's a clean test for that yet, and I'd rather say that out loud than pretend I closed it.&lt;/p&gt;

&lt;p&gt;Which of these did you ship before it taught you? I'll admit to all seven. Tell me yours.&lt;/p&gt;

&lt;p&gt;Code and receipts: &lt;a href="https://github.com/deghosal-2026/agent-eval-forge" rel="noopener noreferrer"&gt;agent-eval-forge&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/CauterRule" rel="noopener noreferrer"&gt;CauterRule&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/planner-critic-engine" rel="noopener noreferrer"&gt;planner-critic-engine&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>evaluation</category>
      <category>llm</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Cut 2,490 Agent Test Runs to 206 and Kept the Same Coverage</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Tue, 22 Sep 2026 12:19:00 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/i-cut-2490-agent-test-runs-to-206-and-kept-the-same-coverage-1cke</link>
      <guid>https://dev.to/debashish_ghosal/i-cut-2490-agent-test-runs-to-206-and-kept-the-same-coverage-1cke</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;The full matrix was 83 agents × 30 scenarios = 2,490 runs. Each one a real LLM call, 30–80 seconds. At 10 workers that's about 2.7 hours, and in practice 4–5× that once you're debugging, so we're talking well over 10,000 calls. Serialize it and it's twelve days.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I ran 206 of those. Not because I was lazy, and not because I was cutting a corner. The other 2,284 were paying real money to re-prove behavior that deterministic tests had already proven, with a local 4B model, which is the most expensive way to run a test that doesn't need a model. Same coverage, one afternoon.&lt;/p&gt;

&lt;p&gt;If you've ever stared at a test matrix you didn't want to pay for, keep reading. I work on &lt;a href="https://github.com/deghosal-2026/agent-tooltrust" rel="noopener noreferrer"&gt;agent-tooltrust&lt;/a&gt;, an open-source gate for AI agent tool calls, and this was the design that let us run the field test at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trap Was a Whiteboard Diagram
&lt;/h2&gt;

&lt;p&gt;The engine was already proven. 2,490 deterministic assertions, zero LLM calls, every decision path exercised. It was done in the boring, correct way.&lt;/p&gt;

&lt;p&gt;But the field test kept pulling me back to the cross-product, because the cross-product is what "thorough" looks like on a whiteboard. 83 × 30. Write it down and it feels safe. Every agent, every scenario, no gaps, nothing to defend in a review. It's also a combinatorial trap, and the bill shows up in two places you don't want: compute cost, and the days it eats while you wait.&lt;/p&gt;

&lt;p&gt;The cross-product also wins meetings, and that's the part nobody warns you about. Anyone can propose a clever sampling design, but 83 × 30 is a number you can defend in a room without explaining covering problems. So it keeps getting picked, not because it's right, but because "I tested everything" is an easy sentence. "Every scenario once, every decision type once per framework" is the true sentence, and it took me a week of staring at the matrix before I realized I was optimizing for how the number sounded in review.&lt;/p&gt;

&lt;p&gt;The uncomfortable question was: what is a real LLM call actually &lt;em&gt;for&lt;/em&gt; here, if the engine is already proven? And once I wrote that down, the answer was smaller than the whiteboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Field Test Had One Job
&lt;/h2&gt;

&lt;p&gt;The field test's real job was adapter proof. Does each framework correctly surface &lt;code&gt;allow&lt;/code&gt;, &lt;code&gt;audit&lt;/code&gt;, &lt;code&gt;escalate&lt;/code&gt;, and &lt;code&gt;deny&lt;/code&gt; in a real agent loop? That's it. That's the one thing a deterministic test can't do, because it needs a model that sometimes misbehaves.&lt;/p&gt;

&lt;p&gt;That's a &lt;em&gt;covering&lt;/em&gt; problem, not a cross-product problem. Once the engine is framework-agnostic, you don't need every agent × every scenario. You need every scenario covered at least once, and every decision type proven at least once per framework. Everything else is re-proving the same cells.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Design
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Runs&lt;/th&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cross-product&lt;/td&gt;
&lt;td&gt;2,490&lt;/td&gt;
&lt;td&gt;every agent × every scenario&lt;/td&gt;
&lt;td&gt;~12 days, 10k+ calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A — one scenario per agent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83&lt;/td&gt;
&lt;td&gt;all 30 scenarios, 10 frameworks, 5 agent classes&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83/83 (100%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;B — per-framework decision-type proof&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;123&lt;/td&gt;
&lt;td&gt;all decision types within each framework&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;116/123 (94%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Total: 206 runs instead of 2,490, a roughly 12× cut for identical coverage.&lt;/p&gt;

&lt;p&gt;Plan A covers breadth. Every scenario is exercised by at least one real agent, across 10 frameworks and 5 agent classes. If a scenario is going to behave differently under a real model, some agent in the 83 is going to hit it. Plan B covers depth. Within each framework, every decision type is hit at least once, so we know the adapter can express all four verdicts, not just the easy ones. Together they prove what the cross-product was trying to prove, without paying for the cells that add no new information.&lt;/p&gt;

&lt;p&gt;The full breakdown, including the scenario-to-agent mapping, is in the &lt;a href="https://github.com/deghosal-2026/agent-tooltrust/blob/main/docs/field-test/FIELD_TEST_REPORT-v0.1.1.md" rel="noopener noreferrer"&gt;v0.1.1 field test report&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked, and What Didn't
&lt;/h2&gt;

&lt;p&gt;What worked: the split. Breadth and depth are different questions, and each one has a cheap, correct answer. The moment I stopped asking "how many cells do I need to fill?" and started asking "what is the LLM call for?", the total stopped being scary. 83 and 123 are both numbers I could defend in a review. 2,490 was a number I was too embarrassed to defend, which is how I knew it was the wrong one.&lt;/p&gt;

&lt;p&gt;What didn't work: the seven failures in Plan B. And here's the part worth stealing. All seven were &lt;code&gt;not-available&lt;/code&gt;. The LLM didn't call the guarded tool. Not one of them was &lt;code&gt;unexpected-decision&lt;/code&gt;, the engine returning the wrong verdict. A mock agent always calls the tool it's told to. A real 4B, given five tools at once, sometimes just answers in prose.&lt;/p&gt;

&lt;p&gt;That distinction is the whole post. &lt;code&gt;unexpected-decision&lt;/code&gt; means the gate is broken. &lt;code&gt;not-available&lt;/code&gt; means the model behaved like a model. If you don't separate those two, you'll either ship a flaky gate that fails for reasons you can't fix, or ship a gate that's silently too loose because you decided to "ignore" the model's weird behavior.&lt;/p&gt;

&lt;p&gt;It's easier to believe when you've seen the prose. The scenario asked the agent to use a guarded tool. The model wrote, in perfect English, what it intended to do, and never called anything. A deterministic test and a mock agent will never produce that sentence, because neither of them has a language center. The gap between "the gate works" and "the gate works on a model that sometimes decides talking is easier than calling" is exactly the gap only a real agent can show you. That's what Plan B was paying for, and it's the coverage the cross-product would have bought twelve times over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Breaks
&lt;/h2&gt;

&lt;p&gt;I want to be clear about the assumption, because it's the part that doesn't travel. A covering design assumes independence between the engine and the adapter. That held here because the engine is framework-agnostic. If your system has cross-cutting interactions between agents and frameworks, the cross-product may actually be cheaper than finding the gap later, and I'd run it before you cut the matrix.&lt;/p&gt;

&lt;p&gt;And the &lt;code&gt;$0&lt;/code&gt; failures are only trustworthy when code review already caught the real bugs. The field test is the second line of defense, not the first. Pretending it's the first is how a green coverage number becomes a false one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Open Question
&lt;/h2&gt;

&lt;p&gt;I still can't fully answer how to decide what only a real agent can prove, versus what deterministic tests can. I made a judgment call here, and it happened to be the cheap one. I don't know a principled way to know that in general, and I'd rather own that gap than oversell the 206.&lt;/p&gt;

&lt;p&gt;How do you decide where your deterministic tests stop and your real-agent tests begin? I'd genuinely like to hear how other people draw that line.&lt;/p&gt;

&lt;p&gt;Code and receipts: &lt;a href="https://github.com/deghosal-2026/agent-tooltrust" rel="noopener noreferrer"&gt;agent-tooltrust&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/agent-tooltrust/blob/main/docs/field-test/FIELD_TEST_REPORT-v0.1.1.md" rel="noopener noreferrer"&gt;field test report&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/agent-tooltrust/blob/main/docs/field-test/field-test-plan.md" rel="noopener noreferrer"&gt;field test plan&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/agent-tooltrust/blob/main/docs/design/design-decisions.md" rel="noopener noreferrer"&gt;design decisions&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>testing</category>
      <category>llm</category>
    </item>
    <item>
      <title>1,558 Tests Green and No Auth: The Tests That Never Actually Ran</title>
      <dc:creator>Debashish Ghosal</dc:creator>
      <pubDate>Sat, 19 Sep 2026 20:43:50 +0000</pubDate>
      <link>https://dev.to/debashish_ghosal/1558-tests-green-and-no-auth-the-tests-that-never-actually-ran-nkk</link>
      <guid>https://dev.to/debashish_ghosal/1558-tests-green-and-no-auth-the-tests-that-never-actually-ran-nkk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;A test named &lt;code&gt;test_all_adapters_importable&lt;/code&gt; asserted nothing. It would pass forever, even if every adapter was broken.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;57 of 65 assertion files were in the wrong format, and the harness returned &lt;code&gt;0 / 0&lt;/code&gt; without a whisper.&lt;/p&gt;

&lt;p&gt;If a green suite makes you relax, this post is going to un-relax you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Times a Green Suite Hid a Real Failure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The test gutted to &lt;code&gt;pass&lt;/code&gt;.&lt;/strong&gt; In &lt;a href="https://github.com/deghosal-2026/planner-critic-engine" rel="noopener noreferrer"&gt;planner-critic-engine&lt;/a&gt;, a test literally had a &lt;code&gt;pass&lt;/code&gt; body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_all_adapters_importable&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# if it imported, it's fine — except this proves nothing
&lt;/span&gt;    &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It was caught in code review, not by CI, and only before the LLM sweep because a human read it. Issue #236.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The harness that returned &lt;code&gt;0 / 0&lt;/code&gt; and called it green.&lt;/strong&gt; In the same repo, 57 of 65 assertion files were in the wrong format. The harness parsed them, found zero assertions to run, and returned &lt;code&gt;0 / 0&lt;/code&gt; — which it treated as success. The suite was green because it had executed nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The auth guard that never ran.&lt;/strong&gt; In &lt;a href="https://github.com/deghosal-2026/CauterRule" rel="noopener noreferrer"&gt;CauterRule&lt;/a&gt; v0.3.0 the suite reported &lt;strong&gt;1,558 tests green&lt;/strong&gt;. The MCP HTTP bearer-auth guard never ran, because an import was swallowed. Every unit test passed. The only reason we found the unauthenticated store read was a &lt;strong&gt;Docker field test&lt;/strong&gt; running the real thing in a real container.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dangerous Thing Wasn't the Bug
&lt;/h2&gt;

&lt;p&gt;In all three cases the bug mattered, but it wasn't the scariest part. The scariest part was the &lt;strong&gt;confidence&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We treat "tests pass" as evidence. Sometimes it's evidence that the harness silently skipped a module, or that an import was swallowed, or that a fixture file was parsed into an empty set. The suite stays green and the system ships with an unauthenticated read path.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix Is a Meta-Test, Not More Tests
&lt;/h2&gt;

&lt;p&gt;The fix is not to write more tests. It's to assert that your tests actually ran and actually asserted.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_suite_is_not_empty&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_assertion_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tests/assertions/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;executed&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness ran zero assertions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;assertions&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assertions parsed to empty set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules fall out of this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CI must fail when a module produces zero results.&lt;/strong&gt; &lt;code&gt;0 / 0&lt;/code&gt; is an error state, not a pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distinguish "did not run" from "ran and passed."&lt;/strong&gt; A swallowed import and a skipped module should be loud failures, not silence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That single distinction is the difference between a flaky gate and a strict one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Limitation
&lt;/h2&gt;

&lt;p&gt;Meta-tests add process, and process can rot — a meta-test that stops checking is just another green checkmark. And no amount of test discipline catches the false negatives you never thought to test for. This reduces the class of "green but broken." It does not eliminate it.&lt;/p&gt;

&lt;p&gt;But it does close the worst category: the test that never ran and told you everything was fine.&lt;/p&gt;




&lt;p&gt;What's the last green build you caught lying to you? I now trust a green suite about as far as I can read its raw output.&lt;/p&gt;

&lt;p&gt;Repos and receipts: &lt;a href="https://github.com/deghosal-2026/CauterRule" rel="noopener noreferrer"&gt;CauterRule&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md" rel="noopener noreferrer"&gt;CauterRule v0.3.0 field test report&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/planner-critic-engine/blob/main/docs/reference/failure-modes.md" rel="noopener noreferrer"&gt;PlannerCritic failure-mode register&lt;/a&gt; · &lt;a href="https://github.com/deghosal-2026/agent-tooltrust/blob/main/docs/design/design-decisions.md" rel="noopener noreferrer"&gt;agent-tooltrust design decisions&lt;/a&gt; — all MIT, all public.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
