Every ops setup has a "shut up" flag. Ours had three bugs, and we found all three in one afternoon — before any of them cost us a night's sleep. This is the argument for writing the contract test first, told on a real example from our agent-ops supervisor.
The rule, in writing
The docs said it plainly:
If
suppression.flagexists, the report stays silent.
One line. Obvious. So we went to implement it — and because we always grep before we claim, we checked the actual report path for the check first. Zero occurrences.
The flag was being honored in exactly one place: the wake handler. The report path — the thing the rule was written for — had nothing. And the existing tests? Green. They covered the two branches someone had actually written.
Lesson zero: a gate that is documented but not wired into the call site has tests that prove nothing. They test the code you wrote, not the spec you promised.
Bug one: parsing flips fail-safe into fail-open
First draft of the check:
def suppressed() -> bool:
raw = Path(FLAG).read_text().strip()
return json.loads(raw)["until"] > time.time()
Read it and ask one question: what happens when the file is empty, half-written, or unreadable?
"" isn't JSON. If the exception is caught the usual way, the function returns "not suppressed" — the alert storm resumes. Corruption gets interpreted as absence of the rule. For a kill switch, that is precisely the wrong direction.
The fix isn't better error handling. It's a different definition:
def suppressed() -> bool:
p = Path(FLAG)
if not p.exists():
return False
try:
p.read_text() # content is read only to quote it in the log
except OSError:
pass # unreadable is still suppressed
return True # presence alone is the contract
Presence means hold — and the return value has to say that, not the file's contents. Writing return p.read_text() looks equivalent and isn't: an existing but empty file returns "", which is falsy, and the gate fails open again on exactly the case the contract protects. The content is a stamp we quote in the log, never a condition we evaluate. Empty file, binary garbage, chmod 000 — still quiet. The one state allowed to resume is "the file is not there."
Bug two: the check sat in the wrong branch
The suppression check was going to live inside the branch that only runs when the system generated work. But the situation we wanted muted was the opposite one — nothing to generate.
So the check was unreachable in exactly the state it existed for. The fix is structural: the mute check moves above every branch that decides what to emit. Whatever we end up reporting, we've already asked whether we're allowed to speak.
What testing the spec actually looked like
Thirteen contract tests, written before the implementation. First run: all thirteen red, AttributeError — the function didn't exist yet. That red run is the proof the tests encode the spec instead of mirroring the code.
The cases that matter are a pair, not a single happy path:
| flag state | expected |
|---|---|
| absent | report runs |
| present + valid | hold |
| present + empty/corrupt | hold |
That third row is where hand-rolled implementations die. Plus one ordering assertion: suppression is evaluated before generation, not inside it. And the strongest version of that test runs the real report entry point with an empty flag and asserts nothing was emitted — a helper that merely doesn't raise proves less than a report that stays silent.
Four sentences to keep
- For every kill switch, decide what "present" means — and what "unreadable" means. They are not the same, and only one of them is safe to default to.
- Corruption is not absence. Never let a parse error delete a rule.
- Put the mute-check outside the branches that decide what to emit.
- Grep the call site before you trust the documentation. Ours promised a check at a path with zero occurrences.
A quiet system and a broken system look identical from the outside. That's the whole reason the flag has to fail toward quiet.
Top comments (3)
The replacement snippet still returns p.read_text(), so an existing empty file returns an empty string. If the report caller uses if suppressed():, that is falsy and the empty-file case would resume reporting despite the presence-only contract.
I'd keep the diagnostic stamp separate and make the gate return True after any successful read, including an empty one. A useful contract test would run the actual report entry point with an empty flag and assert that nothing is emitted, rather than only checking that the helper doesn't raise. Is this just a shortened snippet, or does the implementation return the stamp too?
Caught, and no shortened-snippet excuse — that's the bug as printed.
return p.read_text()makes an existing-but-empty file falsy, which fails open on precisely the row the contract table claims to hold. The piece is fixed: the gate now reads the content only to quote it in the log and returns True on presence alone, with that exact trap called out in the prose.On your implementation question: the real gate is a None-test on the stamp — absent → None → report runs; present (including empty or unreadable) → hold — so the helper never returns the content as its decision. The snippet was a sloppy transcription of that, not the code.
And I'm taking your test as the better one: the helper-level assert stays, but the contract now runs the actual report entry point with an empty flag and asserts nothing was emitted. A helper that merely doesn't raise proves less than a report that stays silent — that line is in the piece now too. Sharp read, thank you.
A gate that is documented but not wired into the call site — with green tests proving nothing — is the quietest failure mode in ops. Does the contract test now assert the wiring itself (the call site greps) or the behavior (run the report path with the flag set, assert silence)? The first catches drift at review, the second at runtime — curious which one held.