DEV Community

Mahiro Hirakawa
Mahiro Hirakawa

Posted on Edited on

Our parser read 0 of 887 lines from our own production data. The gate didn't call it a failure.

A parser we shipped for our own decision log read zero of the 887 lines in the live copy. Not most of them. Zero.

The gate that runs it did not report that as a pass, and it did not report it as a fail either. It reported a third thing, and that third thing is the part worth writing down.

CHAIN parsed=0/887  verdict=UNKNOWN  reason=CHAIN_NOT_RECOMPUTABLE
Enter fullscreen mode Exit fullscreen mode

The setup

We keep a decision ledger. Every ruling a session makes gets an append-only entry, machine-parseable, so a gate can check things like "does every report on disk get declared somewhere" without a human re-reading the whole history. The parser that reads that ledger got rewritten to widen what it could extract.

The new parser passed its test suite. Then someone ran it against the actual live ledger instead of the fixtures.

Why the tests didn't catch it

The suite tested the parser against small, hand-written examples of the grammar. Every one of those examples used the grammar correctly, because a person had written them to demonstrate the grammar. The live ledger had drifted: some real accumulated formatting difference the fixtures never captured. The new parser's assumptions about the input shape were narrower than what 887 real lines actually looked like.

input the parser saw lines who wrote it parsed
fixtures 40 a person demonstrating the grammar 40
the live ledger 887 four months of sessions 0

Same failure shape as testing a filter against clean data and shipping it against dirty data, except here the dirty data is the actual thing being parsed. Nobody had run the new code against it until someone thought to.

What the gate did with 0/887

This is the part I want to hold up.

The check that depends on parsing the ledger, verifying that the causal chain between decisions is recomputable, got a denominator of 887 and a numerator of 0. There are two obvious things to report and both are wrong:

verdict what it claims why it is wrong here
PASS nothing failed true, and vacuous: nothing was examined
FAIL the chain is broken unproven: the chain was never read
UNKNOWN the instrument could not look the only one that is actually true

UNKNOWN is not a hedge. It is a specific claim: this instrument cannot currently tell you PASS or FAIL, and that is different information from either one.

A gate that collapses "I could not check" into "it failed" produces false alarms indistinguishable from real ones, and everyone eventually starts ignoring the noisy gate. Collapsing it into "it passed" is worse. That is a silent green wearing the costume of a real result, and it is exactly the shape that would have kept a 0/887 parser rate invisible if anything downstream had trusted a green it had not looked behind.

The distinction only pays for itself if UNKNOWN is rare enough to notice. If everything is UNKNOWN, nothing is informative either. That failure mode exists too, in the other direction, and we do not have a clean answer for the threshold where a system that reports honest uncertainty stops being useful because it reports it too often.

The other miss, same day: measuring the wrong dimension

A separate piece of the same engine got reorganized into layers, on the working assumption that this would cost time. More indirection, more hops between layers, presumably slower. We measured it: an operation that reads by byte offset against one that walks and greps.

read by byte offset   62 ms
walk and grep          2 ms
                     ----
                      31x
Enter fullscreen mode Exit fullscreen mode

31x is a real number and it is not the number anyone expected to explain. The byte-offset path was not slow because of layering. It was slow because it read the entire body of every record before checking whether the record was even relevant, where the grep path could reject most records without ever reading their bodies.

  • hypothesis going in: the new layering will cost time
  • finding: the layering costs nothing
  • actual defect: one read pattern inside one layer reads more than it needs to

Confirming the wrong half of a hypothesis and refuting the right half in the same measurement is a useful reminder that "did it get slower" and "why" are two different questions. Answering the first one tells you where to look, not what you will find there.

The boring failure that matters more than either

In the same pass, a registry file that is supposed to declare every report a work session produces was checked against what actually exists on disk:

$ ls reports/ | wc -l
21
$ grep -c '^report:' registry.toml
16
Enter fullscreen mode Exit fullscreen mode

Twenty-one report files existed, sixteen were declared, five were orphans. Real work, done, with no pointer to it anywhere a later reader or a gate would look.

Nothing dramatic happened because of this, and it is the failure mode most likely to happen again, because it produces no error, no red test, no crash. Just work that quietly stopped being discoverable. It was caught only because someone thought to compare a list against a filesystem instead of trusting that the list was current.

Where this leaves the parser

Rewritten to handle the shape the live 887 lines actually have, verified against the live copy rather than only the fixtures this time, and the fixtures got the drifted shape added to them so the next rewrite has a chance of catching it before shipping.

The gate still reports UNKNOWN wherever it genuinely does not know, rather than picking the more comfortable of the two wrong answers.

TraceFold, Rust, Apache-2.0, alpha. The parser in this story is not in the public tree yet. This was the DB layer that feeds it, still in motion.

Top comments (0)