DEV Community

Aditya Mishra
Aditya Mishra

Posted on

A silent failure costs you $2,500–8,000 and 6–15 unbillable hours. Here's how to stop finding out three weeks late.

Those numbers aren't mine. They're from a consultant in this community who's cleaned up after them. He also said the part that stings: that diagnosis time is unbillable. You eat it.

Three threads here this month, same shape:

28 consecutive failures running three days before anyone looked. A workflow scoring every lead as zero for six weeks. An agent telling a customer "I can see you were charged" after both lookups came back 403 — run finished COMPLETED.

None of them were red. That's the whole problem. Your error workflow fires on exceptions. These throw nothing.

What you're actually paying for right now

Time to notice: 48 hours to 3 weeks. Usually a client tells you.
Time to diagnose: 6–15 hours digging through executions, unbillable.
Cost of the incident itself: $2,500–8,000.

Per incident. Four a year and that's $10,000–32,000 plus 60 hours you can't charge for.

What I built

Matrix takes what the run claimed — "email sent", "record updated" — and reads the authoritative system to check whether it actually happened. Not the execution log, which the run wrote itself. The actual mailbox.

Three verdicts: contradicted, confirmed, inconclusive. Evidence attached to each.

Two of the three checks need no access to anything:

The claim was made and no tool was ever called — the absence is the evidence.
A lookup returned 403 and the run then stated what that lookup would have shown. The call completed; the result was a refusal. Most wrappers collapse those.

The third needs a connected mailbox: the tool was called, returned 200, the record isn't there.

Why inconclusive matters

If it can't prove the trace was complete, or a tool has a name it doesn't recognise, it refuses to judge rather than calling your working workflow a liar. A monitor that resolves its own uncertainty optimistically is worth nothing on the day it matters.

What it doesn't do

Gmail is the only external adapter. It can't catch a correct call with a wrong argument — right recipient format, wrong recipient, real send, Gmail agrees. And a false claim in text that was never sent has no outcome to read back against.

Setup

One line pasted into Cursor or Claude Code and it wires itself. Or four calls by hand, TypeScript or Python.

matrixverify.dev — free, 20+ users, and I'd rather it got broken than ignored.

Happy to run it against one of your executions and tell you what it finds, including if it finds nothing.

Top comments (2)

Collapse
 
ricart_juncadella_d62f385 profile image
Ricart Juncadella •

For the 200-but-no-record mailbox check, how does Matrix distinguish delayed visibility from a contradicted claim? An immediate read-back could still be inconclusive if the authoritative system hasn’t exposed the record yet.

Collapse
 
aditya_mishra_2417 profile image
Aditya Mishra •

Two things handle it, and neither is a guess.

A verification window with a grace period, not an instant read. Nothing is verified until the provider has had time to index — a run that finished seconds ago isn't checked yet, it's returned as too_recent_to_verify and picked up on a later pass. Guessing early manufactures contradictions, which is the one failure I can't afford.

And the search itself has to be provably complete before a negative means anything. Gmail's result page caps at 25, so if I'm looking for 30 sends and find exactly 25, that's result_page_capped — inconclusive, not a shortfall. Same for the window: if I search ten minutes and the claim covers an hour, a partial result proves nothing. Every collector returns what it found plus the bound it was operating under, and found-at-bound is inconclusive by construction, before any comparison runs.

The rule underneath both: a negative is only authoritative if the question reached the authoritative system and the search space was complete. Otherwise the honest answer is that I don't know.

Where it's still weak: the grace period is a fixed threshold rather than per-provider. A system with a slow async pipeline would need longer, and I'd currently return inconclusive on something that was simply still settling.