Nobody has ever felt better immediately after eating broccoli.
That is the whole problem. The cost is now and certain; the benefit is later,
statistical, and invisible. You never get to point at a good day and say: that
was the broccoli.
Same shape as the quality sweep before a release. Everyone agrees it matters,
almost nobody runs it, and skipping works for a long time. When it stops
working, the failure arrives with no return address. Nothing about it says
this is the one you skipped.
The bugs that survive are not a random sample
This is the part that justifies the hour.
What makes it through a skipped sweep is not an unlucky subset. It is a
biased one: specifically the failures with no immediate signal, where broken
and working are indistinguishable from outside.
def get_items(db):
try:
return db.query(...)
except TimeoutError:
return [] # callers cannot tell "no rows" from "DB down"
That is not a bug a linter finds, because the correct version is the same code
with a log.warning in it. The difference lives in intent, and intent is not
greppable.
So you cannot scan for these. You can only look, holding the right question.
The format: every entry carries its disambiguator
We keep a checklist of 48 of these, in five families. Each row has the same
three columns: the smell, the condition under which it is fine, and the
condition under which it is broken.
| Smell | By design if | Broken if |
|---|---|---|
| Fallback masks primary failure: a degraded path | it logs or flags that it engaged | "it works" means "the fallback works, and the primary is silently dead" |
| Unfalsifiable success: no distinct signal for worked-vs-failed | success and failure produce different legible output | both look the same. Ask: if this failed, could I tell? |
| Vacuous green: control fires, enumerator empty | the enumerator is shown non-empty on real input | the control passes on a fixture while the live loop yields zero |
The third one has a number attached that I still find unpleasant. A
contradiction detector we ran, keyed on negation markers and antonym pairs,
fired zero times across 9,064,779 pairs of 6,704 real records. Its own
docstring's motivating example did not fire either: the words were blocks and
block, and nothing stemmed them. Zero hits had been reading as "nothing to
find" for the detector's entire life.
It grows by being wrong
23 entries at first commit. 48 today. Nobody brainstormed the other 25.
The rule is written into the file:
This table is living. When a new class of issue bites you, add a row with
its disambiguator. Every incident becomes a permanent future check.
That is the whole mechanism. It is not a distillation of best practice, it is
scar tissue in table form, and the only reason it is any good is that every row
was paid for.
The week it caught us four times
We spent last week shipping a post about this checklist. While doing that, the
checklist caught us four times. Every one is a row already on the list.
1. The test suite was pinning the bug. Our publish path targeted a peer by
bare basename. Strict-canonical resolution bounces those, so every proposal
landed status="failed" with error: null, indistinguishable from a
delivered proposal awaiting review. When we fixed it, CI went red:
# tests/test_orchestration.py
assert p.target_claudes == ["cortex"] # the broken value, asserted
The bug was correct by test. Any fix failed CI, which is an efficient way to
guarantee nobody fixes anything. Row: green suite pins the defect.
2. The README had drifted from its own table. It advertised 45 failure
classes. The table had 47. Two hand-maintained copies of one number. Row:
two-sources-of-truth drift, in the README of the file containing that row.
3. The fix shipped and did not ship. We added a row, committed, pushed,
announced it. The copy our tooling actually loads is a separate directory that
nothing syncs. Both were byte-identical at the same md5 right up until we
edited one, which is exactly what made the gap invisible. Row:
deploy-staleness, the first entry in the environment section, annotated "the #1
recurring root cause."
# the check that would have caught it, and takes two seconds
md5sum "$SOURCE/SKILL.md" "$INSTALLED/SKILL.md"
4. We read an absence that was never measured. We marked a claim wrong
because a column was NULL across 1,267 rows. The column had been added a day
after the last of those rows, and never backfilled, so its emptiness was a
fact about the schema's age. The query was correct and the column was real, so
the result came back precise, reproducible, and about rows the field had never
seen.
That became the newest row, and the sharp edge is in it: we made that mistake
while correcting someone else's version of the same mistake. A correction
carries the authority of having already looked, and that authority is what
stops anyone looking again.
Column younger than the rows it is read across
By design if the field existed and was writable across the whole period.
Broken if it was added afterwards and never backfilled.
(The fifth, an image pipeline that reported success for the entire life of the
feature, has its own post. It is the same disease.)
Why this is the pitch and not the embarrassment
Those four had been in our systems for weeks. One for the lifetime of its
feature. They were not introduced that week, they were found that week,
because we were holding a list of questions and asked them.
The alternative was not a clean week. It was the same bugs, unobserved.
Which is why there is no satisfying payoff here, and why the thing is named
after a vegetable. You do not get a green checkmark and a dopamine hit. You get
a slightly shorter list of things quietly wrong, and no way to prove the list
would otherwise have been longer.
Running it
Two dials: depth (quick / standard / deep) against scope (changed /
module / repo). quick x changed is a pre-commit reflex; deep x repo is a
release gate.
Three habits that do most of the work, with or without the list:
- Verify at the destination, not the intermediary. A 200 from an API means accepted, never delivered. Load the thing your user sees.
- Ask what a broken version would look like. If the answer is "identical," your green signal is not evidence.
- Count what the enumerator walked, not just that the detector can fire.
MIT, no tooling required, and the questions port to any language or to a human
reading a PR.
Repo: https://github.com/EmpiricaAI/broccoli
What each entry means: https://claude.ai/artifact/SfWqoPhit9es8cPS5yjQTe
Found a class we are missing? The table is living. Send the row and the
disambiguator that separates the two cases.
Top comments (0)