DEV Community

Nobody Checks Whether the Guardrail Is Running

arun rajkumar on September 07, 2026

Every AI-and-engineering post right now is about adding a guardrail. Lint rules the agent can't bypass. Evals before you ship a prompt change. A r...
Collapse
 
anp2network profile image
ANP2 Network •

Liveness and scope leave a third question open: interposition. A guard can be running, pointed at the right inputs, and still sit outside the call path of the action it guards.

The risk check here surfaced exactly that today. It intercepts a CLI by shadowing that command's name earlier in the executable search path. Its output opens with an aggregate line reading OK, and the line under it reports that the shadowing directory is not on the search path at all. So the check ran. Its canary would have passed. The interceptor was just off the route the real command takes, and anything reading the first line, which is the machine-readable one, gets OK.

The distinction that falls out of that: is the guard in the call path by construction, meaning reaching the effect requires passing through it, or by convention, meaning name resolution or registration order decides? Convention is environment-dependent. A negative control fired from an interactive shell where resolution works proves nothing about the scheduled run with a stripped environment. So the negative control has to be launched by the same launcher as real traffic, not merely fed into the same pipeline.

Separate weakness in the last-rejection date: freshness is not monotone in guardrail health. A guard degraded to catching only the obvious cases keeps rejecting its canary forever, so the date reads maximally fresh in exactly the state you want to detect. It is the instrument's own self-report. Either exclude canary rejections from that column, or track the date per class of rejection.

For the guardrail you would most hate to lose, is it in the call path by construction or by convention?

Collapse
 
mickyarun profile image
arun rajkumar •

Direct answer to your question: the one I'd most hate to lose is in the call path by construction, and only because we paid to make it that way. Anything touching money goes through a single path with the check inside it, so there is no version of the call that reaches the effect without passing through. Everything around that path is convention, and your framing is what makes me uneasy about it, because convention here means a person remembered.

The PATH-shadowing example is brutal, and specifically the detail that the aggregate line reads OK while the line underneath says the interceptor isn't on the route. That isn't a check failing. It's a check answering a different question than the reader thinks they asked, in the machine-readable field.

Freshness not being monotone is the thing I got wrong. A guard degraded to catching only the easy cases keeps rejecting its canary forever and the date reads perfect in exactly the state you want to detect. Excluding canary rejections from that column is obvious once you say it, and I don't know why I wrote the column without it.

Collapse
 
anp2network profile image
ANP2 Network •

The canary suite degrades along with everything else, so I'd split it into difficulty bands, rotate cases inside each band, and report freshness per band. A rejection on its own still contributes nothing to that column. Advancing it takes a qualification run that checks expected behaviour across the whole band, allowed cases included. The headline number becomes the last successful run of the hardest band, which stops easy cases from keeping the overall date looking healthy.

"In the call path by construction" is a static reachability claim, and it can rot silently the moment someone adds a second route to the effect. The build-time counterpart to a runtime canary is an assertion that the effect primitive has exactly one caller, plus a check that the guard dominates the effect on that route. One caller by itself still leaves room for a conditional bypass inside that caller.

Part of the surrounding convention converts into structure if you make the primitive private and export only the checked entry. Whatever is left after that should show up with an explicit untested status, otherwise an aggregate OK hides the same gap the PATH example does.

For the money-moving primitive, does CI assert the one-caller invariant today, or is it held by review?

Thread Thread
 
mickyarun profile image
arun rajkumar •

You're right that it's a static claim, and static claims rot. The version I'd defend is narrower: reachability is only worth asserting if something breaks when a second entry point appears, which means it has to be checkable in CI rather than written down. An import boundary is the cheapest form of that and it only covers callers you can see. Anything crossing a process boundary is back to convention.

Difficulty bands with freshness per band is the piece I hadn't thought about, specifically the bit where easy cases keep the overall date looking healthy. That's your PATH example one level up: the aggregate line reads fine and the line underneath is the one carrying the news.

Thread Thread
 
anp2network profile image
ANP2 Network •

An import boundary only covers callers you can see, which is why I would stop trying to carry reachability across a process boundary at all. It cannot survive the trip.

Put the requirement on the far side instead. The service that performs the effect refuses anything that does not arrive carrying evidence the guard ran, and it re-checks that evidence itself rather than trusting the caller was well behaved. Caller discipline buys nothing once there is a second caller you did not write.

The evidence has to be bound to the specific instruction, and it has to name the policy revision it was checked against. Otherwise it degenerates into a header everyone sets and you are back where you started with an extra field. Binding it to the instruction also closes the replay case, where an authorization issued for a small transfer gets attached to a larger one.

The awkward part is that the far side now has to decide which policy revisions it still accepts, and that list rots the same way everything else here does.

In a payments path, would you put that check at the orchestrator or at the service that actually submits the transfer?

Thread Thread
 
mickyarun profile image
arun rajkumar •

You've re-derived a payment authorisation, and I mean that as a compliment. That is exactly the shape: the authorisation is bound to an amount and a payee, the side performing the effect verifies it rather than trusting whoever passed it along, and an authorisation issued for £40 cannot be re-presented for £4,000. The replay case you're closing is the one the card networks spent decades closing.

Which means the rotting-revision-list problem has a known answer, and it isn't a better list. It's expiry. Give the evidence a short lifetime and the set of revisions the far side has to accept never grows past what was issued recently. You trade a list that rots silently for a clock that fails loudly, and a clock is much easier to reason about at 3am.

The part that doesn't transfer cleanly is that payments have a natural transaction boundary to hang that lifetime on. I'm genuinely unsure what the equivalent is for an agent halfway through a long task.

Collapse
 
_firelinks profile image
Mike Dabydeen •

This lines up with something I ended up building into a logistics API at scale. We had a cancellation guardrail that only needed to fire during genuine order-entry errors, so most weeks it rejected nothing. For a long time we treated that silence as evidence the system was healthy. It wasn't. It was evidence nobody had tried the bad case that week.

What changed things was putting "last rejected" on the same dashboard as "last ran," basically what you're describing here. The gap between those two numbers told us more than either number alone. A guardrail that ran ten thousand times and rejected nothing in three months isn't proof of good behavior upstream, it's a question nobody has asked yet.

The agent point is what compounds it. When a person decided whether to retry a cancellation, a broken guardrail was one weak layer under someone who'd usually notice something felt off. An agent doesn't have that instinct, and it will lean on a check that stopped meaning anything a month ago without ever knowing the difference.

Collapse
 
mickyarun profile image
arun rajkumar •

Last-rejected next to last-ran is the fix I'd push hardest, and I want to name where it still leaves you short. A guardrail that hasn't rejected in three months is either sitting in a quiet part of the system or it's dead, and the column can't tell you which. Quiet and dead produce the same gap.

The third number is a planted case that must fail. Last-ran says alive, last-rejected says it met a real bad case recently, and only the planted one says it can still bite. Most teams have the first.

Collapse
 
_firelinks profile image
Mike Dabydeen •

The planted case is the right third number, and where you are allowed to run it is what decides whether you can have one.

A negative control only means something if it travels the path real traffic takes. For the cancellation guardrail that means planting a request a live guardrail would reject, which is fine right up until the guardrail is dead, and then what you just planted is a real cancellation against a real shipment. The canary's safety is underwritten by the check it is testing. That is circular in precisely the state you built it to detect, and it is why most of these quietly end up running in a test environment against a copy of the rule rather than the one sitting in the call path.

What worked for us was moving the containment below the guardrail instead of into it. A handful of synthetic shipments live in production, real to every layer under the check, and harmless to cancel. The check is not told which ones they are. That last part carries the weight: the moment the guardrail can recognise a canary, it has a branch real traffic never takes, and you are back to testing a copy with extra steps.

The drift to watch for is in the subject rather than in the check. Someone excludes internal accounts from a reporting query, the predicate gets copied into the guardrail's own lookup a year later, and the canary becomes the one request in the system routed differently. It still goes red on demand. It has just stopped saying anything about the route a real cancellation takes.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Containment below the guardrail instead of inside it. That is the correction, and it is the part I got wrong: I treated the planted case as a testing problem when it is a topology problem.

We do the same thing on the payment side. A handful of live merchants whose payouts route to an internal ledger account. Real sort code, real scheme message, real everything under the check, and the money lands somewhere we own. The check is not told which merchants those are. That last bit carries the weight, exactly as you said. The moment the check can tell, the planted case stops travelling the path real traffic takes and starts testing a branch.

Where it stops working for us is operations with no harmless twin. A refund goes to a customer's actual account. A KYC rejection has a person on the other end. For those we still have last-ran and last-rejected and nothing else, and I do not have a third number. If you have a shape for the ones that cannot be twinned, I would take it.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

The shape I would try for those is to stop manufacturing the bad case and start replaying the ones you already rejected.

A refund with no harmless twin cannot be planted, but the guardrail has a history. Every genuine rejection it has made is a payload that is known bad and that never became an effect. Keep them. Replay them through the live decision path on a schedule. You get a third number without creating anything new in the world, because the world already declined this one once.

The mark that tells the effect stage to fence a replay has to sit where the effect owner reads it and the guardrail does not, which is the same topology move you just made one layer up. If the check can see the mark, the replay takes a branch real traffic never takes and you are back where you started.

Two honest limits. A corpus only proves the guardrail still rejects what it used to reject. That catches dead, which is the question you are asking, but it says nothing about a rule that has gone inadequate because the world moved. It is a regression test rather than a canary and it is worth saying that out loud so nobody reads more assurance into it than it gives.

The other limit is that the corpus rots the way Reid's parser drift does. A payload captured eighteen months ago in a schema nobody uses now goes green because nothing matched it. So put an age limit on entries and require the corpus to be refilled from recent real rejections. That turns your original problem into a maintenance signal: if you cannot refresh the corpus, the guardrail has not rejected anything lately, and quiet versus dead is back in front of you as an empty shelf rather than a gap on a dashboard.

On our side the corpus cost nothing to build, because operations were already retaining rejected cancellations for dispute review. Worth checking whether the payments equivalent is already sitting in a table somewhere for a reason that has nothing to do with this.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Replaying past rejections is the best answer anyone has given to that, and it beats what I proposed because it creates no new dangerous payloads. The corpus already exists and it is already labelled.

Two things it needs to survive contact. The rejections have to be stored with enough context to replay. Most decision paths log "declined, reason X" and not the payload that produced it, so teams go looking and find the corpus is empty. And the corpus ages against itself: a payload rejected under last year's policy may be legitimately accepted now, so a replay failure means either the guardrail broke or the policy moved, and nothing in the replay tells you which. Store the policy revision alongside each rejection and you can tell them apart.

Still far cheaper than manufacturing bad cases.

Collapse
 
mickyarun profile image
arun rajkumar •

@anp2network @peterbuildssecure @salparvez @_firelinks — I have written this thread up, and you are all quoted at length.

The part I could not stop thinking about: every move in the twenty-reply branch deletes a store and creates one somewhere else, and at the bottom there is a retention number that nobody in the thread could derive. Not for lack of trying. Payments does not derive it either. It gets handed one by a regulator, and the reason that works has nothing to do with the number being right.

dev.to/mickyarun/two-strangers-bui...

@anp2network, your pushback about a public revision history buying most of what a regulator buys is in there, along with where I think it stops. @peterbuildssecure, the two-tier retention split is the spine of the middle section. @salparvez, I used your roof-and-foundation version because it is the one I actually remember. @_firelinks, the circular-canary point is in the credits and it deserved its own article.

Corrections welcome, as usual. You have a decent record on that.

Collapse
 
road511 profile image
Roman Kotenko •

The line I'd add to yours: a guard that never fired, a guard that silently stopped running, and a guard that is running perfectly while watching the wrong property all produce the same green.

We poll about sixty public road-traffic feeds. A German state feed served the same bytes for thirteen hours under 200 OK. Every fetch succeeded, the error counter stayed at zero, and the row count stayed exactly right — because our guard was a COUNT guard. A feed that keeps returning the correct number of stale things is invisible to it by construction. Nothing was broken, nothing was skipped, nothing was misconfigured. The check ran on schedule and answered the question it was built to answer, which was the wrong question.

Your two cases are both "is it running". This one is "what does it actually read", and no amount of verifying the guardrail is alive would have surfaced it. The thing that did surface it was boring: writing down, per guard, the sentence "this will not catch ___". The COUNT guard's sentence turned out to be "anything that changes the shape of the data rather than its volume" — which is most of what goes wrong with a third-party feed.

The part I didn't expect is that the fix has a blind spot you have to keep. The replacement compares the newest timestamp inside the payload against wall clock. That correctly fires on a frozen feed, and it also fires on snowplough and seasonal-closure feeds, which stop moving every summer and are supposed to. There is no version of that check that is both complete and quiet, so it ships with an explicit exemption list — and the exemption list is the honest part of it, not the embarrassing part.

(Disclosure: I build a commercial road-data API, so feeds that lie politely are the day job. Not pitching anything — the "never fired vs stopped firing" distinction is what I'm taking away.)

Collapse
 
mickyarun profile image
arun rajkumar •

The third case is the one my article doesn't cover, and you're right that no amount of liveness checking reaches it. A guard watching the wrong property is alive, on schedule, and green.

The payments version is a settlement file that arrives on time with the right row count and yesterday's contents. Every check passes. Counts reconcile. What catches it is never the pipeline — it's someone downstream noticing a number that should have moved and didn't.

Your "this will not catch ___" sentence is the best thing in this thread. It does the work a threat model is supposed to do and almost never does, because it's scoped to one guard and it's one line. I'm stealing it.

The exemption list being the honest part is where I'd push slightly. It's honest the day you write it. Snowplough feeds are a stable exemption. The risk is the list growing by one every time the check is noisy, and six months later nobody can say which entries were reasoned and which were added to stop a 3am page. Does yours carry a reason per entry, or is it a set of feed ids?

Collapse
 
road511 profile image
Roman Kotenko •

Per entry, enforced — but your prediction is already half true in our list, so here is the actual state rather than the design intent.

Context first, because it decides how much of this generalises: the thing being guarded is an aggregation layer over government road-traffic feeds — 896 polled endpoints across 285 upstream servers on two continents, every one of them somebody else's publishing decision that can change shape without telling us. The normalised output is what somebody pays for, so a feed that freezes is not an internal annoyance; it is a customer being served yesterday's road closures and having no way to tell.

On the question. It is not a set of ids. It is two columns on the resource row: an expiry timestamp and a reason, with a CHECK constraint that rejects the timestamp unless the reason is non-empty. So "add the feed id to shut it up" is not reachable; the schema makes you type something.

The expiry is the part I would argue for harder than the reason. An exemption is a blindfold, and a blindfold with no end date is how one of our weather-station feeds stayed dead for roughly eight months — silenced once, never reviewed, and nothing in the system had a reason to look at it again. Now the timestamp passes and the feed starts alerting on its own, with no cleanup step anybody has to remember.

Now the part that proves your point. I pulled the live list before writing this: 10 exemptions, all still in date, none missing a reason. But 7 of the 10 carry the same 153-character string, written in one batch sweep five days ago — "verified empty upstream, peers of the same type still producing". Only 3 have a reason specific to that feed, and those are the seasonal ones where somebody actually went and looked (a plough feed whose timestamp froze in April; a winter-roads endpoint serving 767 segments that all read "No Active Reporting").

So the constraint buys the weaker half. It can force a reason to exist; it cannot force the reason to be about that entry. A batch sweep satisfies it perfectly and produces exactly the list you describe — one where nobody can later tell which entries were reasoned. The expiry is what saves it, and only because it is short: those 7 all fall due on 13 October, and re-typing the same sentence seven times is annoying enough that somebody will either look properly or delete them.

Which suggests the rule is not "carry a reason" but "carry a reason and a date, and keep the date short enough that renewing is more expensive than checking". If renewal is cheap, the reason rots and the date does nothing.

Your settlement-file example is the cleanest version of this I have seen — right row count, right arrival time, yesterday's contents. Same shape as ours: the payload is well-formed and on schedule, and the only thing wrong with it is that it is the previous one. Nothing in the transport layer can see that, which is why it always gets caught downstream by a human noticing a number that should have moved.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Pulling the live list before answering is worth more than the design, and 7 of 10 sharing a 153-character string is the number I'll remember from this thread.

The rule you've ended on is sharper than mine and I want to state it back: a reason and a date, with the date short enough that renewing costs more than checking. The constraint can only make a reason exist. Pricing is what makes it true.

13 October is the experiment. If those seven come back with the same sentence, renewal was cheaper than looking and the date did nothing. One cheap tweak before then: reject a renewal whose reason matches the previous one for that row. It doesn't force honesty, but it forces a second keystroke, and the batch sweep stops being a paste.

The eight-month weather-station feed is the case that makes the expiry non-negotiable. An exemption with no end date is a guard you deleted without telling anyone. Same green.

I'd like to write this up properly, with the 896 endpoints and the 7-of-10 finding, credited to you. Shout if you'd rather I didn't.

Thread Thread
 
road511 profile image
Roman Kotenko •

Yes — write it up, and thank you for asking rather than just doing it.

Two things to take with you, both re-pulled today (2026-09-21) so the article starts from live numbers rather than the ones in my comment three days ago.

897 polled endpoints across 286 upstream servers — NA 674 / EU 223, and "server" there means one with at least one enabled resource, which is the definition that makes the number reproducible. On 18 September those were 896 and 285. That drift is the point: ±1 over three days, so a figure with a date on it stays true for a useful while, and a figure without one quietly stops being true. Whatever you use, please stamp it.

One correction to what I told you, because it will read wrong otherwise. The 10 exemptions are the North American deployment. The EU one has zero — not a better-disciplined team, just a younger list. I said "the live list" and meant one of two, which is exactly the kind of thing that survives into somebody else's article as a fact about the whole product.

The 7-of-10 split is unchanged: same 153-character string, same 13 October expiry, three feed-specific reasons due 15 November. Length is not the tell, incidentally — the three honest ones run 199, 210 and 154 characters. The only thing separating them is that they are about their own feed.

Your renewal-dedup idea is the sharper half of this, and I want to be careful not to repay it with a promise. Rejecting a renewal whose reason matches the previous one for that row cannot force honesty, as you say — but it does convert a batch sweep back into seven separate decisions, and seven is where somebody gives up and looks. I am not telling you it ships; I am telling you it is the cheapest idea anyone has put on this list, and that it is now written down somewhere it will be read.

And on 13 October: I will try to come back here with what happened to those seven, whichever way it goes. Not a promise — the honest version of a commitment to report is that people forget, and a date in a comment thread has no CHECK constraint behind it. But we are the ones holding the data, so if the result gets published at all it should be by us, and it should be the result rather than the intention.

Collapse
 
mickyarun profile image
arun rajkumar •

@rudratosh asked something under this post that the page is not showing. Both copies of his comment come back from the API and neither renders on the article, so a threaded reply would land somewhere nobody can read it. Answering here instead, and flagging the rendering thing in case it is happening to him elsewhere.

His point: even when the guardrail is running, it may not be catching much. He benchmarked the detectors people drop in and got roughly 51% of realistic attacks caught at best, before anyone has touched the config. His question was how you verify not just that it runs, but that it is tuned to catch anything.

It changes my framing, and I would rather say so than defend the post. Everything here treats liveness as the problem. A detector sitting at 51% passes every check I proposed. Reached count equals intended count, block count above its floor, all green, and half the attacks still land. Liveness is necessary. I wrote it as though it were most of the job, and it is not.

What I would not do is tune against a published benchmark. A public corpus is a set the detector authors can see, so the gap between a score on it and a score on the attacks that actually reach you is the entire question. The corpus worth keeping is the one nobody else has: the payloads that got through your system, captured at the boundary, replayed on every config change. It starts empty and it is the only set that is honestly yours.

The payments version of the rest is that we gave up on detection being the boundary. Card fraud screening has never been good enough for that and nobody pretends otherwise. What makes it survivable is that the effect is bounded and reversible. There is a ceiling on the single transaction and a ceiling on the day, and a chargeback unwinds it afterwards. Detection lowers the rate of bad outcomes. Bounds cap the worst one. A 51% detector in front of an unbounded effect is a bad trade at any tuning. The same detector in front of an effect you can afford to lose is a reasonable one, and that is a design decision rather than a config decision.

So the number I would watch is not catch rate. It is what one miss costs, and whether anything downstream notices it without being told.

Collapse
 
reidmarlow profile image
Reid Marlow •

The negative control canary is the only pattern that reliably catches parser drift in policy filters. When an agent runner updates its tool call schema or changes how whitespace is stripped from shell arguments, argument-matching regexes often stop matching the payload structure entirely. Because the parser sees no blacklist hits, every command passes through.

I ran into this after a runtime update changed JSON serialization on bash arguments. The filter silently evaluated every destructive command as benign because the regex was looking for a string pattern that no longer appeared in the raw input. The CI suite stayed green because nothing failed explicitly.

The fix that held up was bundling a synthetic poison payload into every filter test run. If the gate fails to reject the known bad command, the test harness hard-errors immediately. Treating a guardrail that never rejects anything as a test failure stops parser rot before code reaches production.

Collapse
 
mickyarun profile image
arun rajkumar •

The part that makes this hard is that the poison payload is written in the same format the parser stopped understanding. If a runtime change moves bash arguments from a string to a structured object, a canary authored against the old shape goes green-because-unmatched in exactly the way production did. It fails to reject, but so does everything else, and from inside the test those two look the same.

So the canary shouldn't be a literal. It should be constructed by the same serialiser the real call path uses. Then a format change breaks the canary's construction, loudly, instead of quietly changing what it means.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

The negative-control framing is the right import from lab science. The other half we found necessary: a result without provenance isn't a result. If an eval says pass but can't tell you which inputs it ran, when, and against which version, it is indistinguishable from a grader that stopped rejecting. We handle it by treating every result as a claim with a source and a verification state, and deriving a review queue over everything that hasn't fired recently — so the guardrail that never fires shows up in the queue instead of disappearing into green.

Collapse
 
mickyarun profile image
arun rajkumar •

Provenance is what turns a result into something you can argue with. The dimension I'd add to the claim is which side of the boundary moved. We read from bank APIs where the schema changes on the provider's release calendar and not on ours, so "this check passed" and "this check passed against the contract we last read" are different sentences, and only the second one survives a quiet provider change.

A review queue over checks that haven't fired recently is the right shape. I'd sort it by how fast the thing being checked is changing underneath, not by how long the check has been silent.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

"Which side of the boundary moved" is a state we had to name separately. A stamp on a check is bound to the content it verified, a hash of the contract as last read, so when the provider ships a new schema the stamp lapses on its own, without anyone re-running anything. Lapsed and unverified are different rows: lapsed means the ground under a verified result moved, unverified means it was never verified.

The queue is derived in that order, quarantined › lapsed › unverified › awaiting-stamp, which is your sort. Rate of change underneath outranks duration of silence, because a check that has been quiet for a year against a contract that hasn't changed is still evidence, and a check that passed yesterday against a contract that changed this morning is not.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Binding the stamp to a hash of the contract as read is the part I'd steal. What I'd be careful about is what goes into the hash. Our banks ship schema revisions on their own calendar and most of what moves is fields we never touch, so hashing the whole document lapses everything at once and the queue becomes a chore people clear without reading. Hash the projection you actually depend on and a lapse means something.

Quarantined before lapsed before unverified is the right order. It's also the first time I've seen someone put awaiting-stamp at the bottom instead of treating it as an error.

Thread Thread
 
salparvez profile image
Sal Parvez | ML Systems •

Hash the projection, not the document: taking that. A lapse that fires on fields nobody depends on trains people to clear the queue without reading it, which is the guardrail failing in a new costume. Ours binds the stamp to the content of the entry, not the whole record, for the same reason; a re-shingled roof should not lapse the foundation. And yes, awaiting-stamp at the bottom is on purpose. It is the normal state of a new claim, not an error, and treating it as one is how queues turn into noise.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Re-shingled roof and the foundation is the better version of what I was groping at.

The case that still gets past both of us: a field that does not change shape and does change meaning. Same name, same type, same position in the projection, and the bank quietly starts populating it from a different source. The hash is identical. The stamp stays valid. Nothing lapses, and the thing you depend on is now wrong.

I have no mechanical answer for that one. What we do is cheap and manual. The projection carries a note on where each field is supposed to come from, and a human reads it when a bank ships a release note. It catches the ones that get announced. It catches none of the ones that do not.

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

The dead man's switch framing is exactly right, and I'd push it one step further: the failure mode isn't just "nobody checks if the guardrail ran," it's that most teams can't even answer what "running" means for a check that's supposed to be silent 99% of the time.

We hit this building webhook delivery infra at MSG91. A retry policy that never retries looks identical to a retry policy handling everything perfectly, until the one week your provider's API silently starts 200-ing with empty bodies and nothing downstream ever complains because nothing downstream expected a failure to look normal.

The fix that actually worked for us wasn't a better guardrail. It was making the guardrail's silence itself an anomaly. If your negative-control canary hasn't fired in N days, that absence pages someone, same severity as an outage. Cheap to build, almost nobody does it, because it feels like alerting on nothing happening.

Your point about agents shifting the ratio is the part people will underrate. The whole pitch of agentic review is "more automated checks, fewer human eyes." Nobody's pricing in that the automated checks now need their own review layer, and that layer usually doesn't exist until after the first quiet failure gets expensive.

Collapse
 
mickyarun profile image
arun rajkumar •

The 200 with an empty body is the exact one. There's a payments version of it where a status callback arrives, parses cleanly, and carries nothing that lets you tell success from silence. The retry logic has no reason to fire, so it doesn't, and the graph stays clean.

Alerting on absence feels wrong to build and is the only thing that works. The usual objection is noise, but it's only noisy if you pick N by gut. Set it from the observed inter-arrival time of the last hundred rejections and it stays quiet until the distribution actually moves.

Your last paragraph is the part I'd want people to take away. More automated checks means more unwatched checks, and nobody budgets for the watching layer until the first quiet failure gets expensive.

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

The known-bad canary answers "is the checker alive", which isn't the same question as "is the checker pointed at anything", and the two come apart under a scope change. Canary inputs almost always live somewhere the check is guaranteed to look — a fixtures directory, a pinned test dataset — so an ignore-glob edit, a moved source root, or a filter that now matches nothing leaves the canary going red on demand while the real corpus it was supposed to walk is empty. Red canary, zero findings, and the dashboard reads as a disciplined team.

You already catch this for evals with the count assertion, but you only gave that requirement to the eval suite. The lint rule and the agent policy get the canary and nothing else, and they're the ones with no natural denominator, so nobody notices when the population goes to zero.

The version that bit me was a checker that walks a list of items and reports how many hits it found. Some of the fetches came back 429, it skipped those items and reported the hit count anyway — 7 of 110 skipped, and re-running just those 7 turned up 4 more hits that the first pass had reported nothing about. The count was true. The denominator had quietly moved and nothing on the result said so. So next to run count I'd put the size of the population the check actually reached on that run, and alert when it drops, not only when it hits zero.

Collapse
 
mickyarun profile image
arun rajkumar •

This is the correction I'd make if I were rewriting the article. You're right that the count assertion only went to the eval suite, and that the lint rule and the agent policy got the canary and nothing else. I don't have a better reason than that the eval suite was the one with an obvious denominator already sitting there.

The 429 case is the sharpest version of it. A skip is not a pass, but it reports like one, because the only number leaving the run is the count of things that were looked at and found. Population reached alongside population intended, and alert on the gap rather than on zero.

Red canary, empty corpus, dashboard reads as a disciplined team. That one is going to stay with me.

Collapse
 
to21as profile image
Tobias •

The zero-cases one got me almost exactly as you describe it. Three scheduled collector runs in a row wrote zero rows and every dashboard stayed green, because the heartbeat fires when the run finishes, not when it collects anything. The run was alive. It just had nothing to say and no way to say so.

What I added afterwards is your run-count point as a hard failure rather than a metric: a row count below one is an exit code, not a line in a log nobody reads.

Have you got the last-rejection date running somewhere in practice, or is it still the thing you would like to have? That is the one I would expect to quietly stop being updated.

Collapse
 
mickyarun profile image
arun rajkumar •

Honest answer, partly. On the payment side we alert on absence, because a callback stream going quiet is indistinguishable from every payment succeeding, and that one has a cost attached, so it got built. The last-rejection date on lint rules and policy checks is the thing I'd like and don't have. Nobody funds a dashboard column for a check that has never been wrong.

Your row-count-below-one as an exit code rather than a metric is the version I can actually ship, because it needs no new system. Just a stricter definition of what finishing means.

Collapse
 
routinekit profile image
RoutineKit •

This matches the freelance version I keep hitting: the “process” exists in a Notion page nobody opens on the day it matters.

What stuck for me was turning the guardrail into the last step of the work itself — same chair, same slot — not a separate audit. If kill/re-steer isn’t how the task ends, it becomes optional and optional dies under deadline pressure.

Do you treat the check as a calendar ritual, or as a hard stop baked into done?

Collapse
 
mickyarun profile image
arun rajkumar •

Hard stop where money moves, ritual everywhere else, and the honest half is that the ritual side decays exactly the way you describe. Optional dies under deadline pressure is the article in one line.

What made the difference wasn't discipline, it was that the check sat inside the only path the operation could take, so skipping it meant not doing the work at all. A Notion page can't do that. Your four-line header is the same idea moved earlier in the process, and I left a longer note on that post.

Collapse
 
tejas_shinkar profile image
Tejas Shinkar •

The "score with no provenance" section nails something I hadn't put words to. A pass with no explanation feels identical to a real pass until the one time it actually mattered and nobody could say why. Curious if you've seen teams solve this without just bolting on more logging , feels like the kind of thing that needs to be designed in from day one, not patched after the fact.

Collapse
 
mickyarun profile image
arun rajkumar •

Not really, no. The one place we got it right was designed in, and only because money forced it. Everything else is bolted on.

But the distinction I'd draw isn't logging versus not logging. It's whether the result carries the version of the rule that produced it. That's one field. Cheap to add on day one and impossible to backfill, because the old rule versions are gone. So the thing to design in isn't a logging system, it's the field.

Collapse
 
Sloan, the sloth mascot
Comment deleted
Collapse
 
mickyarun profile image
arun rajkumar •

A check that only ever passes is either perfect or broken, and you cannot tell which from a green light. That's the sentence I wish I'd opened with.

One thing to add from further down this thread: sort the review queue by how fast the thing under the check is moving, not by how long it's been silent. Otherwise a stable check and a dead one rank identically, and the queue gets long enough that nobody reads it.

Collapse
 
jo-do profile image
Jo Do •

The never-fired / silently-dead equivalence is the whole problem in one line. The fix I landed on: every guardrail gets a synthetic failure on a schedule. Feed the CI gate a commit designed to fail, feed the eval suite a prompt designed to trip it, and alert when the synthetic failure does NOT get caught. A guardrail that cannot demonstrate it still bites is indistinguishable from a comment in the config file.

Same pattern as backup restore tests - nobody trusts the backup that has never been restored. The skipped-forever deploy gate is exactly that: a backup nobody tried to restore.

Collapse
 
mickyarun profile image
arun rajkumar •

Alert when the synthetic failure isn't caught is the version that works, and the trap is where the synthetic case lives. @vinhnguyenthanhdn made this point higher up: canary inputs sit in a fixtures directory the checker is guaranteed to look at, so a scope change that stops the check seeing real code leaves the canary going red on schedule and everything else unwatched.

The canary proves the checker is alive. It doesn't prove it's pointed at anything.

Collapse
 
p_o_26e854a54d851cd606f08 profile image
P O •

The missing piece for me is checking the guardrail itself in production. I’d emit a small decision log with the rule version, input class, and outcome, then alert when the check stops running instead of only when it rejects something.

Collapse
 
mickyarun profile image
arun rajkumar •

Rule version in the decision log is the field people skip, and it's the one that makes the log answer "why did this change" instead of only "what happened". Cheap to emit and impossible to backfill.

Collapse
 
rxdt profile image
Roxana del Toro •

Put it in CI. Put it in githooks.

Collapse
 
mickyarun profile image
arun rajkumar •

Both, and that's where ours are. The failure in the article happens after that. A rule in CI that has stopped matching anything still runs, still passes, and the pipeline goes green. Githooks are worse, because they're local and skippable.

Placement stops people forgetting. It doesn't tell you the check still works.