<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aditya Mishra</title>
    <description>The latest articles on DEV Community by Aditya Mishra (@aditya_mishra_2417).</description>
    <link>https://dev.to/aditya_mishra_2417</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4121767%2Ff10fe8f9-76dc-4e71-8994-6b6c496c633b.png</url>
      <title>DEV Community: Aditya Mishra</title>
      <link>https://dev.to/aditya_mishra_2417</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aditya_mishra_2417"/>
    <language>en</language>
    <item>
      <title>A silent failure costs you $2,500–8,000 and 6–15 unbillable hours. Here's how to stop finding out three weeks late.</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:27:17 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/a-silent-failure-costs-you-2500-8000-and-6-15-unbillable-hours-heres-how-to-stop-finding-out-20g7</link>
      <guid>https://dev.to/aditya_mishra_2417/a-silent-failure-costs-you-2500-8000-and-6-15-unbillable-hours-heres-how-to-stop-finding-out-20g7</guid>
      <description>&lt;p&gt;Those numbers aren't mine. They're from a consultant in this community who's cleaned up after them. He also said the part that stings: that diagnosis time is unbillable. You eat it.&lt;/p&gt;

&lt;p&gt;Three threads here this month, same shape:&lt;/p&gt;

&lt;p&gt;28 consecutive failures running three days before anyone looked. A workflow scoring every lead as zero for six weeks. An agent telling a customer "I can see you were charged" after both lookups came back 403 — run finished COMPLETED.&lt;/p&gt;

&lt;p&gt;None of them were red. That's the whole problem. Your error workflow fires on exceptions. These throw nothing.&lt;/p&gt;

&lt;p&gt;What you're actually paying for right now&lt;/p&gt;

&lt;p&gt;Time to notice: 48 hours to 3 weeks. Usually a client tells you.&lt;br&gt;
Time to diagnose: 6–15 hours digging through executions, unbillable.&lt;br&gt;
Cost of the incident itself: $2,500–8,000.&lt;/p&gt;

&lt;p&gt;Per incident. Four a year and that's $10,000–32,000 plus 60 hours you can't charge for.&lt;/p&gt;

&lt;p&gt;What I built&lt;/p&gt;

&lt;p&gt;Matrix takes what the run claimed — "email sent", "record updated" — and reads the authoritative system to check whether it actually happened. Not the execution log, which the run wrote itself. The actual mailbox.&lt;/p&gt;

&lt;p&gt;Three verdicts: contradicted, confirmed, inconclusive. Evidence attached to each.&lt;/p&gt;

&lt;p&gt;Two of the three checks need no access to anything:&lt;/p&gt;

&lt;p&gt;The claim was made and no tool was ever called — the absence is the evidence.&lt;br&gt;
A lookup returned 403 and the run then stated what that lookup would have shown. The call completed; the result was a refusal. Most wrappers collapse those.&lt;/p&gt;

&lt;p&gt;The third needs a connected mailbox: the tool was called, returned 200, the record isn't there.&lt;/p&gt;

&lt;p&gt;Why inconclusive matters&lt;/p&gt;

&lt;p&gt;If it can't prove the trace was complete, or a tool has a name it doesn't recognise, it refuses to judge rather than calling your working workflow a liar. A monitor that resolves its own uncertainty optimistically is worth nothing on the day it matters.&lt;/p&gt;

&lt;p&gt;What it doesn't do&lt;/p&gt;

&lt;p&gt;Gmail is the only external adapter. It can't catch a correct call with a wrong argument — right recipient format, wrong recipient, real send, Gmail agrees. And a false claim in text that was never sent has no outcome to read back against.&lt;/p&gt;

&lt;p&gt;Setup&lt;/p&gt;

&lt;p&gt;One line pasted into Cursor or Claude Code and it wires itself. Or four calls by hand, TypeScript or Python.&lt;/p&gt;

&lt;p&gt;matrixverify.dev — free, 20+ users, and I'd rather it got broken than ignored.&lt;/p&gt;

&lt;p&gt;Happy to run it against one of your executions and tell you what it finds, including if it finds nothing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>devops</category>
      <category>security</category>
    </item>
    <item>
      <title>Your execution says success. Your database says nothing changed. Which one do you believe?</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:24:35 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/your-execution-says-success-your-database-says-nothing-changed-which-one-do-you-believe-1p1i</link>
      <guid>https://dev.to/aditya_mishra_2417/your-execution-says-success-your-database-says-nothing-changed-which-one-do-you-believe-1p1i</guid>
      <description>&lt;p&gt;Three threads here this month have circled the same gap: a green execution over work that never happened. The answers are all correct and all the same shape — RETURNING id, an If node, throw. Wire it per step, per workflow.&lt;/p&gt;

&lt;p&gt;I got tired of wiring it per step, so I built the version that runs once.&lt;/p&gt;

&lt;p&gt;What it does&lt;/p&gt;

&lt;p&gt;Takes what the run claimed — "email sent", "order updated" — and goes and reads the authoritative system. Not the execution record, which the run wrote itself. The actual mailbox.&lt;/p&gt;

&lt;p&gt;Three verdicts, never two:&lt;/p&gt;

&lt;p&gt;contradicted — the claim and the world disagree&lt;br&gt;
confirmed — every property in the claim was checked and holds&lt;br&gt;
inconclusive — something couldn't be established, so no judgement&lt;/p&gt;

&lt;p&gt;That third one is the whole design. If the trace can't be shown complete, or a tool has a name it doesn't recognise, or the mailbox can't be identified — it refuses to judge rather than calling a working workflow a liar. A monitor that resolves its own uncertainty optimistically is worth nothing on the day it matters.&lt;/p&gt;

&lt;p&gt;Two of the three checks need no access to anything&lt;/p&gt;

&lt;p&gt;The claim was made and no tool was ever called — the absence is the evidence, no external read needed.&lt;/p&gt;

&lt;p&gt;A lookup returned 403 and the run then stated what that lookup would have shown. The call completed; the result was a refusal. Those are different things, and most wrappers collapse them.&lt;/p&gt;

&lt;p&gt;Only the third needs a connected mailbox: the tool was called, returned 200, and the record isn't there.&lt;/p&gt;

&lt;p&gt;A real one it found&lt;/p&gt;

&lt;p&gt;Someone published a production trace at me. Two order lookups came back 403 — the agent got nothing. The final reply still said "I can see you were charged $29 twice on 3 October." Run finished COMPLETED. Nothing in the execution list looked wrong.&lt;/p&gt;

&lt;p&gt;That trace produced nothing the first time I ran it. Four separate causes. Fixing them took two days and three of the four were things I'd assumed worked.&lt;/p&gt;

&lt;p&gt;What it doesn't do&lt;/p&gt;

&lt;p&gt;Gmail is the only external adapter. It can't catch a correct call with a wrong argument — right recipient format, wrong recipient, real send, Gmail agrees. And a false claim in text that was never sent has no outcome to read back against.&lt;/p&gt;

&lt;p&gt;Setup&lt;/p&gt;

&lt;p&gt;One line pasted into Cursor or Claude Code and it wires itself. Or four calls by hand, TypeScript or Python.&lt;/p&gt;

&lt;p&gt;matrixverify.dev — free, zero users, and I'd rather it got broken than ignored. Happy to run it against one of your executions and tell you what it finds, including if it finds nothing.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>devops</category>
      <category>automation</category>
      <category>agents</category>
    </item>
    <item>
      <title>When a silent failure hit you, what did it actually cost?</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Sat, 26 Sep 2026 04:31:54 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/when-a-silent-failure-hit-you-what-did-it-actually-cost-1bnm</link>
      <guid>https://dev.to/aditya_mishra_2417/when-a-silent-failure-hit-you-what-did-it-actually-cost-1bnm</guid>
      <description>&lt;p&gt;Plenty of threads here about catching runs that report success while nothing happened. I've read most of them. What I can't find anywhere is the number.&lt;/p&gt;

&lt;p&gt;So: when it happened to you —&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How long between it happening and someone noticing?&lt;/li&gt;
&lt;li&gt;What did it cost — refund, lost client, hours of digging, something else?&lt;/li&gt;
&lt;li&gt;Did you end up building something for it, or decide it wasn't worth it?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Asking because someone in another thread said diagnosing this is billable work for a consultancy, so a tool that shortens diagnosis shortens the invoice. If that's true for most people here, that changes what's worth building.&lt;/p&gt;

&lt;p&gt;Real numbers more useful than patterns. Even "it cost nothing, we caught it in an hour" is an answer.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>security</category>
      <category>devops</category>
    </item>
    <item>
      <title>Matrix – Check whether your AI agent actually did what it claimed</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Fri, 25 Sep 2026 04:45:55 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/matrix-check-whether-your-ai-agent-actually-did-what-it-claimed-1doa</link>
      <guid>https://dev.to/aditya_mishra_2417/matrix-check-whether-your-ai-agent-actually-did-what-it-claimed-1doa</guid>
      <description>&lt;p&gt;Agents report success for actions that never happened — no error, clean trace, and every observability tool reads it as a success, because they're all reading the agent's own account of itself.&lt;/p&gt;

&lt;p&gt;This doesn't read the trace differently. It queries the authoritative system instead — the actual Gmail mailbox — and returns confirmed / contradicted / inconclusive with the evidence attached: the account queried, the search window, and what was found in it.&lt;/p&gt;

&lt;p&gt;Two failure classes, different evidence. If there's no tool call in the trace at all, the absence is the evidence and nothing external is needed. If the call happened and returned cleanly but nothing landed, only the mailbox can tell you.&lt;/p&gt;

&lt;p&gt;The parts I'd most want criticised:&lt;/p&gt;

&lt;p&gt;Inconclusive is a first-class verdict. If the trace can't be shown complete, or the tool span has a name I don't recognise, or the mailbox can't be established — it refuses to judge rather than calling a working agent a liar. A false accusation costs more than a missed detection. I may have that balance wrong.&lt;/p&gt;

&lt;p&gt;It cannot catch a correct call with a wrong argument. Asked to mail one person, confidently mails another — the send is real, Gmail confirms it, verdict is confirmed, correctly. The instruction is recorded next to the arguments so a human can see it. No verdict catches it.&lt;/p&gt;

&lt;p&gt;Gmail only so far. LangChain via a callback handler, or the SDK by hand. No n8n, no CrewAI. Zero users — this is day one.&lt;/p&gt;

&lt;p&gt;matrixverify.dev&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>automation</category>
      <category>testing</category>
    </item>
    <item>
      <title>Matrix – Check whether your AI agent actually did what it claimed</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:26:59 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/matrix-check-whether-your-ai-agent-actually-did-what-it-claimed-18pi</link>
      <guid>https://dev.to/aditya_mishra_2417/matrix-check-whether-your-ai-agent-actually-did-what-it-claimed-18pi</guid>
      <description>&lt;p&gt;Agents report success for actions that never happened — no error, clean trace, and every observability tool reads it as a success, because they're all reading the agent's own account of itself.&lt;/p&gt;

&lt;p&gt;This doesn't read the trace differently. It queries the authoritative system instead — the actual Gmail mailbox — and returns confirmed / contradicted / inconclusive with the evidence attached: the account queried, the search window, and what was found in it.&lt;/p&gt;

&lt;p&gt;Two failure classes, different evidence. If there's no tool call in the trace at all, the absence is the evidence and nothing external is needed. If the call happened and returned cleanly but nothing landed, only the mailbox can tell you.&lt;/p&gt;

&lt;p&gt;The parts I'd most want criticised:&lt;/p&gt;

&lt;p&gt;Inconclusive is a first-class verdict. If the trace can't be shown complete, or the tool span has a name I don't recognise, or the mailbox can't be established — it refuses to judge rather than calling a working agent a liar. A false accusation costs more than a missed detection. I may have that balance wrong.&lt;/p&gt;

&lt;p&gt;It cannot catch a correct call with a wrong argument. Asked to mail one person, confidently mails another — the send is real, Gmail confirms it, verdict is confirmed, correctly. The instruction is recorded next to the arguments so a human can see it. No verdict catches it.&lt;/p&gt;

&lt;p&gt;Gmail only so far. LangChain via a callback handler, or the SDK by hand. No n8n, no CrewAI. Zero users — this is day one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>langchain</category>
      <category>devtools</category>
      <category>showdev</category>
    </item>
    <item>
      <title>I Built a Tool That Verifies What AI Agents Actually Did</title>
      <dc:creator>Aditya Mishra</dc:creator>
      <pubDate>Sat, 12 Sep 2026 06:39:47 +0000</pubDate>
      <link>https://dev.to/aditya_mishra_2417/i-built-a-tool-that-verifies-what-ai-agents-actually-did-33jn</link>
      <guid>https://dev.to/aditya_mishra_2417/i-built-a-tool-that-verifies-what-ai-agents-actually-did-33jn</guid>
      <description>&lt;p&gt;AI agents can report success for actions that never happened.&lt;br&gt;
The trace looks clean, but observability tools often rely on the agent's own account of what happened.&lt;br&gt;
So I built Matrix — a tool that checks the real system instead of trusting the agent.&lt;br&gt;
For example, if an agent says it sent an email, Matrix checks the actual Gmail mailbox and returns:&lt;br&gt;
CONFIRMED — it happened&lt;br&gt;
CONTRADICTED — it didn't&lt;br&gt;
INCONCLUSIVE — there isn't enough evidence&lt;br&gt;
The verdict comes with the evidence behind it.&lt;br&gt;
Gmail only for now. Zero users. Day one.&lt;br&gt;
I'm building this because AI agents need a way to prove what they actually did — not just report that they did it.&lt;/p&gt;

&lt;p&gt;Try Matrix:&lt;br&gt;
&lt;a href="https://matrix-snowy-beta.vercel.app%E2%81%A0%EF%BF%BD" rel="noopener noreferrer"&gt;https://matrix-snowy-beta.vercel.app⁠�&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'd love feedback from anyone building AI agents: How are you currently verifying agent actions?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
