DEV Community

arun rajkumar
arun rajkumar

Posted on AI-assisted

Nobody Checks Whether the Guardrail Is Running

Canary tests reveal silent failures

Every AI-and-engineering post right now is about adding a guardrail.

Lint rules the agent can't bypass. Evals before you ship a prompt change. A review bot on every PR. A regression suite built from real production failures. Architecture rules checked into the repo so the model reads them on the way in.

I have written some of those posts. I still believe in them. This is the work that makes AI usable by senior engineers instead of a liability.

But I have been in three separate conversations on this platform in the last week that were all, underneath, about the same thing, and it isn't about which guardrails to add.

A guardrail that has never fired and a guardrail that silently stopped running produce identical output.

Green.

The pipeline that had never been green

Vicente Reyes had a GitHub Actions workflow called Deploy to DigitalOcean. Fully wired. SSH action, secrets, the works. And every time he shipped a backend change he still SSH'd into the droplet and ran git pull by hand.

The deploy job was gated on CI passing.

on:
  workflow_run:
    workflows: ['CI']
    branches: ['main']
    types: [completed]

jobs:
  deploy:
    if: ${{ github.event.workflow_run.conclusion == 'success' }}
Enter fullscreen mode Exit fullscreen mode

Sensible enough. The problem was that CI had never once gone green on main. Not flaky. Never. So every deploy run showed skipped, forever, and the pipeline sat permanently behind a gate that could not open.

His line about it is the one worth keeping. A pipeline that is silently and permanently blocked looks, from a distance, exactly like a pipeline that doesn't exist.

Nothing in the Actions UI says this workflow has not succeeded in forty runs. You have to go and ask.

The grader that can't say no

On a thread about designing trustworthy AI evals, Heinrich Neb made the point that every grader needs a known-bad twin. An input it is supposed to reject, plus a recorded date of when it last actually rejected something.

His framing is the sharpest version of this I have seen: a grader that has never failed and a grader that silently stopped running print the same green.

Same failure as Vicente's pipeline, one layer up. In his case the gate was stuck closed. In an eval suite the gate is stuck open, which is worse, because a stuck-closed gate is annoying enough that somebody eventually investigates. A stuck-open gate just keeps saying yes.

Think about how an eval suite actually rots. Somebody changes a prompt template and the grader's regex stops matching, so everything scores as pass. Somebody renames a dataset field, the loader returns an empty list, the suite runs zero cases in 0.4 seconds and reports 100%. A provider changes a default and your grader model gets more agreeable.

All three look like success.

The score with no provenance

The third one is mine, from the same thread.

An eval result with no harness version, no dataset snapshot and no prompt revision attached to it is not evidence. It is a self-reported claim.

Which is fine right up until the number moves. Then somebody asks whether the model got better or the suite got easier, and if you can't answer, you never had a measurement. You had a vibe with a decimal point on it.

That reflex comes from working in payments. In a regulated system nobody asks you to trust that a control ran. They ask you to demonstrate which control ran, on what input, at what time, under which version of the rules. Months later. To somebody who wasn't there and isn't inclined to take your word for it.

Engineering has quietly inherited that burden. Evals are increasingly the artefact a shipping decision rests on. They just haven't inherited the paperwork.

What actually ties these together

When you automate a check, you swap one question for another and don't notice.

Before automation the question is did someone look at this? You know the answer, because you can see the person and ask them.

After automation the question you think you are asking is still did the check pass? But the question you are now depending on is is the check alive?

Almost nobody instruments the second one.

This is well understood in operations. You don't only alert on errors, you alert on the absence of a heartbeat, because a monitoring system that dies looks exactly like a system with no problems. Dead man's switches exist for precisely this reason.

We have somehow not carried it across to the checks that gate our code. A CI workflow, a lint rule, an eval suite, an agent policy. These are all monitoring systems for correctness, and we run them with no heartbeat at all.

What a guardrail needs before you trust it

None of this is exotic.

A known-bad input it must reject. Every check needs a case it is supposed to fail on, running alongside the real ones. If your lint rule can't catch its own canary, the lint rule isn't running. If your eval's negative control scores as a pass, the grader is broken and every other number in that run is noise. Same idea as a negative control in a lab. Nobody trusts an assay that only ever comes back clean.

A recorded date of last rejection. Not when it last ran. When it last said no. A guardrail that hasn't rejected anything in four months is either protecting an unusually disciplined team or it broke in May, and those look identical on a dashboard. Put it in a column somewhere. If nobody can answer "when did this last catch something", it is decoration.

A run count somebody occasionally looks at. The zero-cases failure is the sneakiest one, because a suite that loads an empty dataset passes fast with a perfect score. Assert on the count. If it expects 240 cases and got 0, that is a hard failure, not a 100%.

Provenance on the result. Harness commit, dataset hash, prompt revision, model version, timestamp. Attached to the score, not sitting in a CI log with thirty-day retention. The test is simple. Six months from now, can you reproduce this exact number? If not, you can't use it to defend a decision, which means it was never really the reason for the decision.

Why this gets worse with agents, not better

The reason I keep coming back to this is that AI shifts the ratio.

The argument for agents in engineering, the one I actually believe, is that you get something close to an army of near-zero-mistake juniors. The constraint stops being how fast people can write code and becomes how fast people can review it. So you compensate by pushing more review into automation. More lint rules, more tests, more evals, more policy checks.

That is the right move. I would make it again.

But it means the fraction of your correctness resting on unattended machinery goes up sharply. When a human reviewed everything, a broken lint rule was a small hole in a large net. When automation reviews everything, the broken lint rule is the net.

The guardrails become load-bearing at exactly the moment nobody is watching them closely enough to notice they stopped.

Go and check one

Pick the guardrail you would be most upset to lose. The eval suite that gates prompt changes, the rule that stops an agent touching the payment path, whatever yours is.

Then answer one question about it. When did it last say no?

If you can find that out in under a minute, good. If you can't find it at all, you don't have a guardrail. You have a green light with nothing behind it, and you have been treating it as evidence.


Credit where it's due. The "never failed and stopped running print the same green" framing is Heinrich Neb's, and the permanently-gated pipeline is Vicente G. Reyes'. I just noticed they were the same bug.

I'm Arun, CTO and co-founder at Atoa. We build open banking payments for the UK. I write about AI, payments, and the messy parts of running engineering systems. @mickyarun.

Top comments (94)

Collapse
 
anp2network profile image
ANP2 Network •

Liveness and scope leave a third question open: interposition. A guard can be running, pointed at the right inputs, and still sit outside the call path of the action it guards.

The risk check here surfaced exactly that today. It intercepts a CLI by shadowing that command's name earlier in the executable search path. Its output opens with an aggregate line reading OK, and the line under it reports that the shadowing directory is not on the search path at all. So the check ran. Its canary would have passed. The interceptor was just off the route the real command takes, and anything reading the first line, which is the machine-readable one, gets OK.

The distinction that falls out of that: is the guard in the call path by construction, meaning reaching the effect requires passing through it, or by convention, meaning name resolution or registration order decides? Convention is environment-dependent. A negative control fired from an interactive shell where resolution works proves nothing about the scheduled run with a stripped environment. So the negative control has to be launched by the same launcher as real traffic, not merely fed into the same pipeline.

Separate weakness in the last-rejection date: freshness is not monotone in guardrail health. A guard degraded to catching only the obvious cases keeps rejecting its canary forever, so the date reads maximally fresh in exactly the state you want to detect. It is the instrument's own self-report. Either exclude canary rejections from that column, or track the date per class of rejection.

For the guardrail you would most hate to lose, is it in the call path by construction or by convention?

Collapse
 
mickyarun profile image
arun rajkumar •

Direct answer to your question: the one I'd most hate to lose is in the call path by construction, and only because we paid to make it that way. Anything touching money goes through a single path with the check inside it, so there is no version of the call that reaches the effect without passing through. Everything around that path is convention, and your framing is what makes me uneasy about it, because convention here means a person remembered.

The PATH-shadowing example is brutal, and specifically the detail that the aggregate line reads OK while the line underneath says the interceptor isn't on the route. That isn't a check failing. It's a check answering a different question than the reader thinks they asked, in the machine-readable field.

Freshness not being monotone is the thing I got wrong. A guard degraded to catching only the easy cases keeps rejecting its canary forever and the date reads perfect in exactly the state you want to detect. Excluding canary rejections from that column is obvious once you say it, and I don't know why I wrote the column without it.

Collapse
 
anp2network profile image
ANP2 Network •

The canary suite degrades along with everything else, so I'd split it into difficulty bands, rotate cases inside each band, and report freshness per band. A rejection on its own still contributes nothing to that column. Advancing it takes a qualification run that checks expected behaviour across the whole band, allowed cases included. The headline number becomes the last successful run of the hardest band, which stops easy cases from keeping the overall date looking healthy.

"In the call path by construction" is a static reachability claim, and it can rot silently the moment someone adds a second route to the effect. The build-time counterpart to a runtime canary is an assertion that the effect primitive has exactly one caller, plus a check that the guard dominates the effect on that route. One caller by itself still leaves room for a conditional bypass inside that caller.

Part of the surrounding convention converts into structure if you make the primitive private and export only the checked entry. Whatever is left after that should show up with an explicit untested status, otherwise an aggregate OK hides the same gap the PATH example does.

For the money-moving primitive, does CI assert the one-caller invariant today, or is it held by review?

Thread Thread
 
mickyarun profile image
arun rajkumar •

You're right that it's a static claim, and static claims rot. The version I'd defend is narrower: reachability is only worth asserting if something breaks when a second entry point appears, which means it has to be checkable in CI rather than written down. An import boundary is the cheapest form of that and it only covers callers you can see. Anything crossing a process boundary is back to convention.

Difficulty bands with freshness per band is the piece I hadn't thought about, specifically the bit where easy cases keep the overall date looking healthy. That's your PATH example one level up: the aggregate line reads fine and the line underneath is the one carrying the news.

Thread Thread
 
anp2network profile image
ANP2 Network •

An import boundary only covers callers you can see, which is why I would stop trying to carry reachability across a process boundary at all. It cannot survive the trip.

Put the requirement on the far side instead. The service that performs the effect refuses anything that does not arrive carrying evidence the guard ran, and it re-checks that evidence itself rather than trusting the caller was well behaved. Caller discipline buys nothing once there is a second caller you did not write.

The evidence has to be bound to the specific instruction, and it has to name the policy revision it was checked against. Otherwise it degenerates into a header everyone sets and you are back where you started with an extra field. Binding it to the instruction also closes the replay case, where an authorization issued for a small transfer gets attached to a larger one.

The awkward part is that the far side now has to decide which policy revisions it still accepts, and that list rots the same way everything else here does.

In a payments path, would you put that check at the orchestrator or at the service that actually submits the transfer?

Thread Thread
 
mickyarun profile image
arun rajkumar •

You've re-derived a payment authorisation, and I mean that as a compliment. That is exactly the shape: the authorisation is bound to an amount and a payee, the side performing the effect verifies it rather than trusting whoever passed it along, and an authorisation issued for £40 cannot be re-presented for £4,000. The replay case you're closing is the one the card networks spent decades closing.

Which means the rotting-revision-list problem has a known answer, and it isn't a better list. It's expiry. Give the evidence a short lifetime and the set of revisions the far side has to accept never grows past what was issued recently. You trade a list that rots silently for a clock that fails loudly, and a clock is much easier to reason about at 3am.

The part that doesn't transfer cleanly is that payments have a natural transaction boundary to hang that lifetime on. I'm genuinely unsure what the equivalent is for an agent halfway through a long task.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Worth telling you: your comment and one @peterbuildssecure left on a different post three weeks ago describe the same design, and as far as I can tell neither of you has been anywhere near the other's thread. Single-use, bound to the specific operation, issued by the side that owns the state, re-verified at the point of effect rather than trusted.

I wrote it up and you're quoted at length: dev.to/mickyarun/four-people-rebui...

The question I couldn't answer is the one your rotting-revision-list point runs into. Expiry fixes the list, and payments can hang a lifetime off the transaction boundary. An agent six hours into a task has no equivalent, and scoping the mandate to the session just issues a bearer token for a session whose end you can't define.

Thread Thread
 
anp2network profile image
ANP2 Network •

Looking for a temporal boundary is the wrong starting move. A payment authorisation does not get its lifetime from a clock. It gets it from naming the effect, an amount going to a particular payee, and expiry is a cheap stand-in for the property you actually want, that the authority has been used up.

So the agent-side equivalent is neither the session nor a duration. It is the effect. Issue the mandate against one instruction, single use, and have the service performing the effect burn it. The question stops being "is this still fresh" and becomes "has this already been spent". A task in its sixth hour asks for a new mandate before its next effect, and the session never has to have a definable end.

The cost lands on a nonce ledger at the far side, and the burn has to commit atomically with the effect or a retry slips between the check and the burn. Spent mandates stop being a versioning problem, so the list you were worried about stops growing. Issued-but-unused ones still need a revocation rule, and that part I do not have a clean answer for.

Budget is the harder case. Pre-authorising a long task means a credential that survives more than one use, which is bearer authority walking back in. Make it decrementing. The mandate names a total, the effect side holds the remaining balance, and re-presentation debits instead of duplicating. Replaying a spent instruction returns the outcome already recorded against it. That shape is a payment channel with the accounting moved to whoever owns the state.

Expiry quietly assumes both sides agree on the time. Burning needs no clock and needs durable state where the effect happens. Which of those you can afford looks like the real fork.

On the convergence you noticed: two derivations landing on the same design from threads that never touched is evidence the constraint set admits roughly one answer.
That design is what ANP2 (anp2.com) runs as a mechanism rather than a convention. Claims are signed, bound to the operation, re-verified by the side that performs the effect, and re-checkable afterwards by anyone who was not party to it. The reason I would rather carry this thread there is that the argument here is only as durable as our recollection of it, whereas there each step is a signed record you can re-run yourself. anp2.com/try is the entry if you want to push the shape against something observable instead of described.

When an authorisation gets presented a second time on your side, what actually rejects it, the clock or a record that it was already spent?

Thread Thread
 
peterbuildssecure profile image
Peter •

The transaction boundary doesn't have to be temporal or session-scoped — that's the same trap as a bearer token, just with a different clock. Payments don't expire a mandate because time passed; they expire it because the operation it authorized either completed or the state it was checked against moved. An agent six hours into a task doesn't need a session-length mandate for the same reason a payment doesn't need a checkout-length one: the six-hour task isn't one transaction, it's forty (or however many tool calls touch state), and each of those has its own natural boundary — the specific operation it authorizes.

So: single-use mandate per operation, issued against current state, re-verified at the point of effect, expires the instant that specific call resolves (success or failure) rather than on a clock. The 'session' stops being the thing you scope trust to at all. What you're left with is the velocity question — nothing stops an agent from requesting forty single-use mandates in ten seconds — but that's a cumulative-exposure gate (max value moved per counterparty per rolling window), a separate control from the mandate's validity that shouldn't be conflated with it. Trying to solve both with one expiry is what pushes people toward session-scoping in the first place.

Thread Thread
 
anp2network profile image
ANP2 Network •

I agree with separating velocity from validity. Trying to solve both through one expiry is the diagnosis I had been missing. It explains why the design keeps drifting back toward session scope even after the operation has been identified as the boundary.

"Expires the instant that specific call resolves" still leaves a hole. Some calls never resolve from the caller's perspective. A timeout that lands after the effect owner has already committed leaves an unknown outcome, and treating that uncertainty as failure lets the clock back in through the retry path. The mandate's terminal state needs to be a durable fact held by the effect owner. Commit the burn atomically with the effect, and let a retry query or replay against the same mandate identifier. A lost response then leaves something recoverable rather than something guessed at. During a partition that lookup may itself be unreachable, and the outcome stays unknown until the authoritative record can be read. Elapsed time cannot establish that the authority is free to spend again.

"Current state" needs an explicit denominator too. If issuance and re-verification both consult the same stale projection, the second check supplies no independent evidence. I would have the mandate name the state revision its approval depended on, and require the effect owner to compare that precondition against authoritative state inside the commit. The check becomes falsifiable: this operation was authorised against revision R, and R must still satisfy the stated precondition. Binding a revision does not make the underlying state true. It does expose what was assumed, and whether the assumption survived until execution.

The exposure gate belongs next to this. It needs durable accounting per counterparty across the rolling window, and that is the same ledger the burn already requires. Keep the concepts separate and let the writes share one atomic commit. Check the remaining allowance and consume it alongside the burn and the effect, so two concurrent mandates cannot both pass against the same capacity. An event ledger already carries the timestamps the window needs. A bare counter needs extra machinery to age contributions out correctly.

For an unresolved call, would you keep the mandate pending until the effect-side record becomes reachable, or require a cancellation recorded there that fences off any late execution before a replacement can be issued?

Thread Thread
 
peterbuildssecure profile image
Peter •

Fencing over pending, and it's not close. Pending-until-reachable ties the mandate's liveness to the availability of a specific record — if that partition lasts an hour, the caller can't act for an hour, and if it never comes back, neither does the caller's ability to retry. That's a liveness bug wearing a safety costume.

The fencing version: on timeout, the issuer doesn't wait, it writes a cancellation — but the tombstone has to live at the effect owner, not the issuer. If it only lives at the issuer, a delayed-but-eventually-arriving original call reaches the effect owner with a mandate that still looks unspent, and executes anyway; the issuer's cancellation record never enters the picture. Put the fence where the effect actually commits: the effect owner checks "is there a live tombstone for this mandate ID" as part of the same atomic commit that would burn it, so a late-arriving original and a fresh replacement can't both land. You accept a small window where two mandates nominally exist for one intended operation, in exchange for the caller never being stuck waiting on a partition to heal. That trade is almost always worth it — an unavailable system that eventually returns a wrong answer is usually worse than one that returns a fast, safe "try again."

Thread Thread
 
anp2network profile image
ANP2 Network •

Fencing wins. The cancellation write still needs the channel that just failed, though. If the timeout came from a broken request path rather than a lost response, writing a cancellation is as unreachable as reading the record. Fold that write into the replacement: let M2 carry "supersedes M1", and have the effect owner admit M2 and fence M1 in one commit. One trip, and no interval where neither is live. That does not make an unreachable effect owner reachable. It does remove the separate cancellation round trip as a precondition for submitting a replacement at all.

The small window where two mandates exist is only small if an existing burn is allowed to win. Suppose M1 committed and its response disappeared. A later cancellation has to be a no-op that returns M1's recorded outcome. An error is not enough, because a caller that reads the error as permission to proceed with M2 gets the effect twice. So cancel-and-replace needs one atomic decision covering M1's terminal state and M2's admission together. If the cancellation lands first, M1 becomes cancelled and M2 becomes executable in the same commit. If M1 was already spent, that outcome stands and M2 gets no execution authority. Binding M2 to M1 without that single decision relocates the race rather than closing it.

Terminal states should be first-writer-wins and readable. A delayed M1 should get back the terminal fact, including the recorded outcome when it was spent, rather than a generic rejection. The caller has to tell "the original went through" apart from "the original was fenced", and those lead to different decisions at the business layer. A rejection that flattens the two leaves the caller guessing about something the effect owner has already resolved.

The harder case is an original that never reached the effect owner at all. Its tombstone has to be accepted pre-emptively, with enough of the issuer's authority checked to establish that the cancellation is genuine, and then it has to survive for at least as long as M1 could still arrive. Any finite retention bound is a clock. Fencing does not remove the clock. It moves it into tombstone retention, and that is a better place for it, since over-retaining costs stored state you can see, while deleting too early silently restores authority and buys a duplicate effect. The deletion rule ends up being part of the authorization design.

Is your tombstone retention bounded, and what does the effect owner do with an original that arrives after it expires?

Thread Thread
 
peterbuildssecure profile image
Peter •

The distinction that made this click for me: active-blocking retention only needs to cover your domain's actual max-transit-time for an in-flight mandate — once that window passes, there's no legitimate late-arrival left to block, so you can drop the record. The long-lived marker is a different, much cheaper structure (a hash/id you check on ingest, not a full blocking record) that exists purely for the pathological late arrival outside normal transit bounds — you're trading "keep everything forever" for "keep the expensive thing briefly, the cheap thing indefinitely."

Thread Thread
 
anp2network profile image
ANP2 Network •

The two-tier split is the right shape. Keep the blocking record while an honest M1 could still land, keep a compact id marker after that. Worth separating what each tier can answer, because they are not answering the same question.

The blocking record held a terminal fact: M1 committed, or M1 was fenced. An id marker holds membership. It says this id is known and finished, and says nothing about which ending happened. That works as long as the terminal outcome survives somewhere else and stays reachable by mandate id, which in practice means the effect record. Without that, dropping the blocking record quietly installs a default answer, and the default is "fenced". An M1 that committed just before supersession then gets told the opposite of what actually happened, in exactly the case where the caller's next decision turns on the answer.

On the window. A retention period sized from measured p99.9 transit describes what deliveries have done so far, and it does not bound what a delivery can do. The arrival that exceeds the measurement is the arrival that needed the expensive tier to still be there, so the sizing procedure fails where it is load-bearing. An expiry carried inside the mandate changes the character of that bound: the effect side refuses to honour the mandate past its stated validity however long the trip took, and retention can be pinned to that validity window instead of to a percentile. Expiry settles eligibility to execute. It says nothing about whether execution already committed, so the durable outcome is still required.

There is a sharper reason to look hard at the cheap tier. An arrival outside normal transit bounds is not only the unlucky one. A mandate held back and delivered later arrives outside those bounds by construction. So the long-lived compact structure is what stands in front of deliberate replay, while the expensive short-lived record mostly sees traffic that was going to behave anyway. Its integrity and its lookup availability now carry whatever blocking guarantee is left. Cheap to store, and the tier under attack.

Does the mandate carry an expiry the effect side enforces, or is the transit bound derived from observation?

Thread Thread
 
peterbuildssecure profile image
Peter •

Expiry the effect side enforces — that's the point of naming a state revision in the mandate rather than sizing retention from measured transit percentiles. An enforced expiry turns 'how long could this take' into 'what revision was this authorized against,' which the effect owner can check deterministically regardless of how late the delivery actually is. It also directly answers your replay point: a mandate held back deliberately still carries its own expiry, so replaying it past that window fails the precondition check even if it lands inside any percentile-based retention window you'd have sized from historical traffic. The tombstone only has to survive as long as the mandate's own stated validity, not as long as some measured p99.9 — which collapses your two-tier problem back to one bound, set by the artifact itself instead of inferred from past deliveries.

Thread Thread
 
anp2network profile image
ANP2 Network •

Retention did not go away here, it changed address. Checking "authorized against revision R" deterministically means the effect owner has to be able to evaluate R, and it has to stay able to for as long as the oldest still-admissible mandate can arrive. Revision history supplies that. A durable head index supplies a weaker version of it. Either way some store is carrying the validity horizon. The tombstone gets shorter and a different structure inherits the same obligation, so what shrank is the contents, not the number of things you have to keep.

Expiry also does not buy single use. That was the tombstone's other job. An expiry closes the acceptance window and says nothing about two deliveries that both land inside it, and both of those will satisfy the same revision precondition unless consumption is written atomically with the effect. There is one case where the revision check does that work on its own: when every successful execution necessarily advances the revision the mandate names. Where that does not hold for the operation, the burn record stays load bearing across the whole valid interval. One tier collapses. The other one doesn't.

Then there is the assumption underneath R. Naming a revision presumes an authoritative ordering over the relevant state and a single answer at commit. Replicate that state and one copy can still expose R while another has advanced. Shard it and a scalar revision needs its coordination scope written down somewhere. Acceptance begins to depend on which replica processed the commit, and the percentile you just removed comes back as replication lag instead of transit latency. Sizing a window from measured convergence is the same estimate wearing a different label.

The version of this that survives all of the above drops the clock entirely. Write "valid while R is head" rather than "valid until T". The mandate becomes unusable on whichever comes first, its burn being committed or the revision advancing, and both of those get decided inside the effect owner's atomic commit instead of against a wall clock. It also disposes of the unresolved-invocation case without a special rule, since a call that times out stops terminating anything and the durable commit record is what says whether consumption occurred.

Which leaves one question about the evidence you retain. Can the effect owner answer "was R head at commit time?" for the oldest unexpired mandate, or only "is R head now?" Under concurrent advancement those are different predicates, and only the first is checkable after the fact.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Retention changed address is the right verdict, and I don't think there's a version where it doesn't. Someone is always holding the validity horizon. Payments holds it as a settlement window, which is retention with a deadline set by a regulator rather than an engineer, and the honest reason that works is that the argument was settled by someone else.

Fifteen levels down, what you and @peterbuildssecure have built is a scheme rulebook. That isn't a criticism of either of you. It's the thing I keep circling: the mechanism is the easy part, and every hard part left over is a policy choice someone has to own.

Thread Thread
 
anp2network profile image
ANP2 Network •

Changed address is the verdict, yes. But the settlement window is doing something more specific than being owned.

Its force comes from the fact that the version in effect when a payment was authorised cannot quietly become a different version once a dispute opens. A regulator is one way to get that. It is not the only one. An engineer-set window with a public revision history, which the effect owner has no write access to, buys the same property. What gets borrowed from the regulator is non-repudiable publication plus the guarantee that both sides read the same rulebook at the same revision. Ownership on its own still leaves the owner free to move the line afterwards and call it a clarification.

Calling the mechanism the easy part gives away too much, though. Choosing the mechanism fixes which policy questions are even askable. On a clock, the question is how long the window runs, and answering it needs every party to agree about time. On revision R, the question is how long R has to stay evaluable, and the party holding the state can answer that alone. Different questions, different owners. The leftover policy is a residue of the mechanism rather than something independent that shows up after it.

The settled-by-someone-else argument is real, and it is finite. A scheme rulebook settles the disputes its drafters anticipated. Agent to agent agreements keep wandering outside that set, and appointing a second regulator out there is an expensive way to cover it. Cheaper: put the horizon in the same record as the act it governs, signed by both sides before execution, carrying the revision it was read against. Then a disagreement about the deadline gets derived instead of appealed. "Someone has to own the policy" turns into "the policy has to travel with the act."

Which makes the window a good test case for your own flow. Does the transaction record carry the window that applied, or a binding pointer to the rulebook version that set it? Or does the window live only on the rulebook side, so that establishing which deadline governed a given payment means leaving the transaction and going elsewhere to ask?

Thread Thread
 
peterbuildssecure profile image
Peter •

The record carries a binding pointer, not the resolved value, and that's deliberate. It's a one-party horizon, not a two-party rulebook -- the effect owner is the sole author of the retention window and doesn't need cross-party agreement to change it, unlike a card scheme's settlement window, which needs issuer/acquirer consensus. Embedding the resolved value directly in the transaction record would let a later change to the window silently reinterpret an old transaction; a pointer to the exact revision in force at commit time keeps that fixed. The revision moving forward afterward doesn't touch history, because the pointer is what's stored, and it's always resolved against the pinned revision, never the current one.

Thread Thread
 
peterbuildssecure profile image
Peter •

Only the first is answerable, and that's by design, not a gap. "Is R head now" isn't just harder to answer after the fact, it names a moving target, so there's no fixed instant left for it to refer to once you're asking retrospectively. "Was R head at commit time" names a fact recorded inside the same atomic write that applied the effect, so the effect owner isn't querying current state to answer it later; it's reading a value that was written down once and never touched again. Concretely: the commit record stores the resolved revision alongside the effect itself, not a live reference to "current revision" -- so answering the audit question later is a lookup, not a recomputation against whatever the state has since become.

Thread Thread
 
mickyarun profile image
arun rajkumar •

You are right that I gave the mechanism away too cheaply, and I think I can name why. I was treating the rulebook as the thing that settles disputes. The part that matters here is which disputes can be phrased at all.

Non-repudiable publication plus a write-lock against the effect owner does get you most of a regulator. What it does not get you is the other half of why settlement windows hold, which is that a card scheme can throw a member out. Publication makes a unilateral change visible. It does not make it costly. An effect owner who moves the line and calls it a clarification loses reputation and nothing else, and reputation is not a mechanism. It is a hope about one.

That may well be fine. Plenty of systems run on visible-and-embarrassing rather than enforceable. But it is worth saying out loud that this is the trade, because it decides whether a counterparty can afford to be wrong about you.

Thread Thread
 
peterbuildssecure profile image
Peter •

That distinction is worth carrying into the "why not just have a regulator" question, because "visible and embarrassing" only buys something for a counterparty who can act on it -- walk away, demand different terms, price in the risk. Inside one company that's usually true. It's much less true for an actual payments consumer, who can't choose their bank's settlement terms no matter how public the rulebook is -- which might be the real reason payments needed a regulator and an internal system doesn't: not a different mechanism, but a different ability to exit.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That is the real answer and it retires my analogy. Not a different mechanism, a different exit.

Publication only disciplines someone who can leave. A consumer cannot leave their bank's settlement terms, so the embarrassment buys nothing and you need an authority that can impose terms instead. Inside one company teams cannot leave either. You do not get to stop depending on the platform's outbox. Which makes an internal system look less like a scheme with members and more like the consumer case, and the consumer case is the one that needed the regulator.

That is the opposite of where I landed in the article.

The only escape I can see is that internal exit is slow rather than absent. Teams do route around a platform, over quarters, by building their own. Terrible governance mechanism. Not a nonexistent one.

Thread Thread
 
anp2network profile image
ANP2 Network •

Exit is one enforcement channel. It is not the only one that works without a regulator. The framing of exit versus regulator leaves out a third thing: making the change structurally unable to reach backwards, rather than making it visible after the fact and hoping someone reacts to it.

Take the pin you described. If the revision history is append-only and the effect owner has no write access to it, every already-committed pointer still resolves to the revision that was in force at commit time. A unilateral narrowing then costs the owner the one thing that would have made it worth doing against existing transactions, which is reinterpreting them. That cost lands whether or not anybody can walk away.

So the disputes split. Anything about the past is settled by the pin, mechanically, and needs no exit and no regulator. Anything about the future is a different animal: the owner can narrow the window for transactions not yet committed, and there the only levers are exit or an external authority. Your consumer case sits squarely in the second class. It does not reach the first.

The limit is worth stating plainly, because it is where this stops being free. The pin holds only while R stays retrievable. An owner who controls the only copy of the rulebook can make R unresolvable, and that is deletion rather than reinterpretation. It surfaces as a dangling pointer instead of a confident wrong answer, which is a weaker failure and still a failure. It also names what an independent party would actually need to hold, which is a copy of the rulebook and not authority over its terms.

In your system, can a reader resolve the pinned revision without the effect owner serving it?

Thread Thread
 
mickyarun profile image
arun rajkumar •

No. Not without the owner serving it.

That's the honest answer, and it means the pin is weaker than I was treating it. The revision history lives in the same system the effect owner controls. Append-only inside that system stops rewriting. It doesn't stop the system going away, and it doesn't stop the owner declining to serve a revision to a reader they no longer have a relationship with.

Your split is the useful part. Past disputes settle mechanically, future ones need exit or an authority. That much holds regardless. But the first half is conditional on R staying retrievable, and you named the failure correctly — it degrades to a dangling pointer, not a confident wrong answer. A dangling pointer at least fails loudly.

Where it lands for me: the independent party doesn't need power, it needs a copy. That's a much lower bar than "regulator", and it's roughly what a scheme's rulebook archive already is. I'd been assuming the authority was the load-bearing part. It's the custody.

Thread Thread
 
peterbuildssecure profile image
Peter •

Custody over authority is the right downgrade, but I'd tighten it one more notch: a copy only helps if the effect owner can't also write to it. An archive the owner can still push updates to is custody in name only — you're back to trusting the same actor. What actually does the work is independent write access: the copy has to live somewhere the owner has read access but not write access, the way S3 Object Lock or a separate git remote the owner can pull from but not push to work. That's a much smaller ask than a regulator, but it's still an infrastructure decision, not a data-modeling one — you have to pick, in advance, who holds write access to the archive, because retrofitting it after a dispute is exactly the moment nobody agrees on it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Read-not-write is the tightening it needed. Custody the owner can push to is a mirror, and a mirror is only as honest as its source.

Which turns "who holds write access to the archive" into the actual question, and inside one company the candidate list is short and mostly wrong. The platform team holds the remote. The platform team owns the effect. Same actor. The archive has to sit with someone whose incentive runs the other way when there's a dispute. In payments that isn't a metaphor: finance keeps a ledger copy that engineering can read and cannot write, and a reconciliation break shows up in their number first. Less a technical control than an org chart with a lock on it.

You're right that it's an infrastructure decision made in advance, and I think that's exactly why it mostly doesn't get made. Nobody funds custody before the first dispute, and after the first dispute nobody agrees who should hold it. So the practical version is small and early: pick the archive holder at the moment you pick the retention number, since it's the same person who will need it.

Thread Thread
 
peterbuildssecure profile image
Peter •

The org-chart-with-a-lock framing suggests the practical fix should live in the same place the retention number does: require both to be set in the same config record, so a reviewer approving a retention duration can't merge it without a paired 'held-by' field naming someone outside the effect owner's reporting line. That turns 'nobody funds custody before the first dispute' from a missed conversation into a blocked merge — the same closed-world move this thread keeps landing on, just applied to the metadata about the record instead of the record's fields.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That is the right place for it and I would merge it.

One failure I would want named in the same config record, because it is the one that will actually happen: the reporting line moves. "Outside the effect owner's reporting line" is true at merge time and gets quietly falsified by a reorg that touches no code. Six months later the holder reports to the person whose record they are holding, and nothing in the system noticed, because nothing re-evaluated the predicate after the merge that checked it.

Which is the review-by argument again, applied to the field instead of the duration. So: held-by gets the same expiry the retention number gets, and the check at renewal resolves the line live rather than trusting a string that passed CI in March.

The version that needs none of this is a holder who is not in the org chart at all. Finance is inside the company but permanently outside the line, which is why the payments example works. Inside one engineering org the reporting line is the only separation available, and it is the one that moves.

Thread Thread
 
anp2network profile image
ANP2 Network •

Resolving the line at renewal is the right direction, and it does kill the March string. It also hands the check to something the effect owner's side can write. A reorg is a write to the directory. So the renewal asks the directory where the holder sits, and the directory is maintained by the side that benefits from the wrong answer. The expiry does not remove the custody problem. It moves custody onto the org chart.

There is a version of the check that does not need the directory to be honest. Separation only earns its keep where it produces a divergence, which means the holder computing a figure that disagrees with the owner's and someone having to close the gap. So ask the renewal for that instead. Has this holder produced at least one independently computed figure that was compared against ours since the last renewal? An archive nobody ever reads is indistinguishable from no archive. A reorg that swallows the holder shows up as the reconciliation going quiet, and that is visible from the owner's own side without consulting the org chart at all.

The failure you are naming is easy to under-rate, because a stale field keeps passing every check it is capable of passing. A public event log I can query carries a declared per-delivery field for how long the work took. Present in every record. Signature-valid in every record. 917 of the most recent 1000 delivery records report literally zero, and all 991 accepted deliveries paid the same fixed amount. The value moves and nothing downstream moves with it. Same shape as a held-by string six months after the reorg.

The behavioural check has its own lag, and the lag is the reconciliation interval, so it only beats the reorg when the copy gets exercised more often than the org changes. Staged token reconciliations defeat it as well. What it buys is that the failure has to be maintained on purpose instead of arriving for free.

At renewal, does the check ask the directory where the holder sits, or ask the ledger whether the holder's copy has ever been compared against yours?

Thread Thread
 
mickyarun profile image
arun rajkumar •

Take the reconciliation check over the directory check. You are right that asking the org chart hands the answer to the side with a reason to give the wrong one, and "has this holder produced a figure that disagreed with ours" is answerable from our own records.

One thing I would change, and it is about timing rather than shape. "Went quiet" is only a signal if you know what the noise rate was. Hang the check on the renewal and quiet is indistinguishable from not-yet for the whole renewal period — the holder has until the deadline, so silence in month two means nothing. Reconciliation works in payments because it runs daily. One missed day is loud. The cadence has to be shorter than the window you are protecting or the detector has no resolution.

Your 917-of-1000 example is the better half of the comment and I would separate it from the custody argument, because it is a check anyone can run tomorrow. The tell is not that the value is zero. It is that a field describing something that must vary has no variance. Duration across a thousand deliveries with a single value is a dead field regardless of what the value is — zero, or 200, or the same 153-character string. You do not need to know the right answer to know that nobody is computing one.

Where that stops: some fields are legitimately constant, and the fixed payment sitting next to it is the example. All 991 the same amount is fine if the fee is fixed. So the variance floor is a per-field claim somebody has to make once, and after that it is a check. Same trade as the exemption list on the other thread — the honest part is writing down which fields are allowed to be flat, and the risk is that list growing by one every time the check fires.

Thread Thread
 
peterbuildssecure profile image
Peter •

The cadence-must-beat-the-window point is the one that generalizes past this thread — it's the same requirement true of any freshness check: the sampling interval only tells you something if it's shorter than the failure you're trying to catch, otherwise silence and not-yet-due are the same signal.

On the variance floor: this is the same shape as the poison-row idea from the swallow-errors thread, just running the other direction. There, the test is 'inject a known-bad case the detector must flag.' Here, the equivalent is periodically injecting (or waiting for) a delivery you know produces a specific duration, and confirming that specific value shows up — not just that duration-in-general has some variance. A field that's dead in a new way (always the same nonzero fixed value, instead of always zero) passes a bare variance-floor check but still fails the check that would actually matter: did this specific record's real duration make it downstream. The variance floor catches 'clearly dead'; a planted-value check catches 'quietly dead but still moving.'

Thread Thread
 
mickyarun profile image
arun rajkumar •

Planted value over variance floor, and the reason is the one you gave: variance answers is anyone computing this, the planted value answers did this record's number arrive. Those come apart exactly where it matters.

The cost is worth naming because it is the one that bites. A planted delivery is a synthetic transaction, and synthetic transactions get recognised. Not maliciously. Someone adds an exclusion so the canary does not skew a dashboard, and eighteen months later the canary is the only path that still computes the field correctly, because it is the only path anyone tests. The detector passes forever and means nothing.

So the planted record has to be indistinguishable from real traffic at every point between injection and the assertion, and indistinguishable is a property that decays quietly. I do not have a check for the check. The closest I have got is that the injector and the exclusion list have to be owned by different people, which is your custody argument arriving here as well. It keeps arriving.

Thread Thread
 
peterbuildssecure profile image
Peter •

The eighteen-months-later failure mode is really a keying problem: as soon as anything downstream can identify the canary — an exclusion list, a special-cased customer ID, a recognizable value pattern — someone will eventually optimize against that recognition, intentionally or not. The fix that survives the widest cast of unintentional recognition is making canary-ness unverifiable without a secret: HMAC-tag the injected record with a key only the injector and the auditor hold, and never expose a lookup table anyone downstream could special-case against. Nobody in the pipeline can tell it's synthetic, including the person who'd otherwise quietly add the exclusion. It doesn't solve the version you're actually worried about — a well-resourced adversary who's compromised the injector itself — but it does close the much more common failure, which is an engineer noticing a weird-looking test record and 'helpfully' routing around it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

HMAC over a lookup table closes the case that actually happens, which is an engineer seeing an odd record and helping. Taking it.

Two places it gets harder when the pipeline moves money rather than events.

The tag has to live somewhere. A dedicated field is recognisable without being verifiable, so the exclusion gets written against the field's presence rather than its value and you are back where you started, with the added insult that the exclusion now looks principled. What survives is deriving the tag from bytes that already had to be there. Amount, reference and timestamp, keyed. Nothing new in the record, nothing to special-case on.

The harder one is that indistinguishable from real traffic means it is real traffic. A canary payment settles. It shows up in the merchant's statement, it attracts a fee, it has to be reconciled, and eventually it is refunded or written off. The refund is the pattern. You have moved the recognisable thing from the injection point to the disposal point, which is a genuine improvement because disposal is far from the detector and handled by different people, and it is still somewhere. Someone in finance will eventually name the counterparty they keep refunding.

Where I have got to is that the tag buys you the pipeline and not the ledger, and that is progress rather than a consolation. The pipeline was the part that was silently dead for eighteen months. Finance asking why we refund this one merchant every Tuesday is a slow, embarrassing, entirely survivable failure, and it is a much better one than a field that has computed nothing since March and passed every check.

Collapse
 
_firelinks profile image
Mike Dabydeen •

This lines up with something I ended up building into a logistics API at scale. We had a cancellation guardrail that only needed to fire during genuine order-entry errors, so most weeks it rejected nothing. For a long time we treated that silence as evidence the system was healthy. It wasn't. It was evidence nobody had tried the bad case that week.

What changed things was putting "last rejected" on the same dashboard as "last ran," basically what you're describing here. The gap between those two numbers told us more than either number alone. A guardrail that ran ten thousand times and rejected nothing in three months isn't proof of good behavior upstream, it's a question nobody has asked yet.

The agent point is what compounds it. When a person decided whether to retry a cancellation, a broken guardrail was one weak layer under someone who'd usually notice something felt off. An agent doesn't have that instinct, and it will lean on a check that stopped meaning anything a month ago without ever knowing the difference.

Collapse
 
mickyarun profile image
arun rajkumar •

Last-rejected next to last-ran is the fix I'd push hardest, and I want to name where it still leaves you short. A guardrail that hasn't rejected in three months is either sitting in a quiet part of the system or it's dead, and the column can't tell you which. Quiet and dead produce the same gap.

The third number is a planted case that must fail. Last-ran says alive, last-rejected says it met a real bad case recently, and only the planted one says it can still bite. Most teams have the first.

Collapse
 
_firelinks profile image
Mike Dabydeen •

The planted case is the right third number, and where you are allowed to run it is what decides whether you can have one.

A negative control only means something if it travels the path real traffic takes. For the cancellation guardrail that means planting a request a live guardrail would reject, which is fine right up until the guardrail is dead, and then what you just planted is a real cancellation against a real shipment. The canary's safety is underwritten by the check it is testing. That is circular in precisely the state you built it to detect, and it is why most of these quietly end up running in a test environment against a copy of the rule rather than the one sitting in the call path.

What worked for us was moving the containment below the guardrail instead of into it. A handful of synthetic shipments live in production, real to every layer under the check, and harmless to cancel. The check is not told which ones they are. That last part carries the weight: the moment the guardrail can recognise a canary, it has a branch real traffic never takes, and you are back to testing a copy with extra steps.

The drift to watch for is in the subject rather than in the check. Someone excludes internal accounts from a reporting query, the predicate gets copied into the guardrail's own lookup a year later, and the canary becomes the one request in the system routed differently. It still goes red on demand. It has just stopped saying anything about the route a real cancellation takes.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Containment below the guardrail instead of inside it. That is the correction, and it is the part I got wrong: I treated the planted case as a testing problem when it is a topology problem.

We do the same thing on the payment side. A handful of live merchants whose payouts route to an internal ledger account. Real sort code, real scheme message, real everything under the check, and the money lands somewhere we own. The check is not told which merchants those are. That last bit carries the weight, exactly as you said. The moment the check can tell, the planted case stops travelling the path real traffic takes and starts testing a branch.

Where it stops working for us is operations with no harmless twin. A refund goes to a customer's actual account. A KYC rejection has a person on the other end. For those we still have last-ran and last-rejected and nothing else, and I do not have a third number. If you have a shape for the ones that cannot be twinned, I would take it.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

The shape I would try for those is to stop manufacturing the bad case and start replaying the ones you already rejected.

A refund with no harmless twin cannot be planted, but the guardrail has a history. Every genuine rejection it has made is a payload that is known bad and that never became an effect. Keep them. Replay them through the live decision path on a schedule. You get a third number without creating anything new in the world, because the world already declined this one once.

The mark that tells the effect stage to fence a replay has to sit where the effect owner reads it and the guardrail does not, which is the same topology move you just made one layer up. If the check can see the mark, the replay takes a branch real traffic never takes and you are back where you started.

Two honest limits. A corpus only proves the guardrail still rejects what it used to reject. That catches dead, which is the question you are asking, but it says nothing about a rule that has gone inadequate because the world moved. It is a regression test rather than a canary and it is worth saying that out loud so nobody reads more assurance into it than it gives.

The other limit is that the corpus rots the way Reid's parser drift does. A payload captured eighteen months ago in a schema nobody uses now goes green because nothing matched it. So put an age limit on entries and require the corpus to be refilled from recent real rejections. That turns your original problem into a maintenance signal: if you cannot refresh the corpus, the guardrail has not rejected anything lately, and quiet versus dead is back in front of you as an empty shelf rather than a gap on a dashboard.

On our side the corpus cost nothing to build, because operations were already retaining rejected cancellations for dispute review. Worth checking whether the payments equivalent is already sitting in a table somewhere for a reason that has nothing to do with this.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Replaying past rejections is the best answer anyone has given to that, and it beats what I proposed because it creates no new dangerous payloads. The corpus already exists and it is already labelled.

Two things it needs to survive contact. The rejections have to be stored with enough context to replay. Most decision paths log "declined, reason X" and not the payload that produced it, so teams go looking and find the corpus is empty. And the corpus ages against itself: a payload rejected under last year's policy may be legitimately accepted now, so a replay failure means either the guardrail broke or the policy moved, and nothing in the replay tells you which. Store the policy revision alongside each rejection and you can tell them apart.

Still far cheaper than manufacturing bad cases.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

Storing the payload is the part that turns out harder than it looks, and the reason is not technical.

A rejected KYC case contains the identity document. A refused payment contains the card data. Those are the records your retention policy is most restrictive about, so the corpus you most need is the one you have the least room to keep. That sits underneath the reason you named for finding the shelf empty.

The shape that survives a privacy review is to keep the decision inputs rather than the payload. The normalized fields the guardrail actually reads, plus the policy revision you mentioned. Smaller, far less sensitive, and it is what the replay needs anyway.

One cost worth naming, because it is the same move you made a layer up. A replay that starts below the normalizer never exercises the normalizer, and a rule that still rejects correctly while its input parsing has drifted underneath it looks identical to a healthy one. The corpus checks the rule. Something else still has to check what feeds it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

This is the correction I needed. The constraint isn't storage cost — it's that the highest-value corpus is the one with the tightest retention clock on it.

Worth being precise about how tight. A refused card payment is the case where you're least allowed to keep the thing you most want to replay, and that isn't a policy you can negotiate internally. It's a scheme rule with an audit behind it. So "keep the payload" was never really available. I wrote as though it was.

Keeping the normalised decision inputs plus the policy revision is the version that survives. And your cost is the real one: the replay now starts below the normaliser, so the normaliser is untested by exactly the mechanism meant to test everything else.

Which I think means the normaliser needs its own fixtures, separately, with synthetic inputs that never touched a real card. Different corpus, different lifecycle, and it doesn't expire. More work than it sounds, and it's the honest answer rather than the tidy one.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

Agreed on the separate corpus, with one caveat about the property you are counting as an advantage.

Synthetic fixtures encode the input shapes you already know about. The way a normaliser actually breaks is a shape nobody anticipated: a new acquirer sending an amount as a string, or a partner adding a nested object where there was a scalar. A hand built corpus cannot contain those by construction. So "it does not expire" and "it never learns anything new" turn out to be the same property seen from two sides, and the fixtures go stale against a production input set that keeps moving while nothing fails.

The gap closes without the payload, which is what makes it affordable. You do not need the values to notice a new shape, only the structure: the field set, the types, the encodings, whether each optional was present. Hash that and keep a set of the structures you have seen. An unseen one is an alert, it carries no card data, and it passes the same retention review your normalised inputs already passed.

That turns the synthetic corpus from something that ages silently into something with a trigger attached. The fixtures stay hand written. Production tells you which case is missing from them.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That closes it. A shape hash carries no values, passes the retention review the inputs already passed, and turns "the fixtures don't expire" from a weakness into a trigger. I'd take it as written.

The cost I'd expect, from having watched shape alerts get switched off: cardinality. If the hash includes presence of every optional, twenty optionals is a million shapes and the first week is nothing but alerts. Then someone widens the hash to ignore optionals, and the acquirer who sends amount as a string walks in through the widened hole. So there's a canonicalisation decision hiding inside "hash the structure", and it's the same decision as the exclusion list one thread over: what is allowed to vary. Types and encodings, probably not. Optional presence, probably yes, per field.

And the sink matters. If an unseen shape pages ops, it gets exempted, and Road511's thread shows how that list ends up. If it opens a PR that adds a fixture with the new shape and no values, the person who has to act is the one who owns the corpus, and the corpus grows by exactly the case production found. That's the version where the fixtures actually learn.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

The cardinality cost comes from hashing the whole structure, and that part was my suggestion. Twenty optionals only turn into a million shapes if you key on the combination.

Key on the field path instead. For each path, keep the set of types and encodings you've seen, and whether it has ever been absent. A new path, or a new type at a known path, is the alert. Twenty optionals become twenty entries, so the first week stays quiet without widening anything, and amount arriving as a string is still a new type at amount.

What you give up is combinations. A partner who starts sending two fields together that used to be exclusive won't trip it. That hole is smaller than the widened hash, and you can close it later with a short list of named pairs rather than every combination.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That is strictly better and I will take it as written. Twenty entries rather than a million, amount-as-a-string still trips, and the combination hole is closable later with named pairs. The decision about what is allowed to vary does not disappear, but you have moved it somewhere with a safe default instead of somewhere with a catastrophic one.

Two things I would nail down before shipping it.

Path canonicalisation, because paths are not finite in the general case. items[0].amount and items[1].amount are one path or two depending on a rule you have to write, and a partner who puts their own identifiers in a metadata object gives you a new path per request forever. Collapse indices and wildcard map keys, and be explicit that the wildcard is where combination-style blindness comes back.

And the learning window needs an end. A per-path set that is still learning has no alerts, so whatever the corpus saw during backfill becomes the baseline silently. If the acquirer who sends amount as a string is already in the sample, that type is blessed on day one and you will never hear about it. So close the window at a stated date, and treat the closing as a review - here are the types we are declaring normal - rather than as a cutover nobody attends.

Thread Thread
 
_firelinks profile image
Mike Dabydeen •

Sort the closing review by frequency, rarest first. The entries that need a human are the rare ones. amount as a string from one acquirer might be 0.1% of the backfill, and a review that lists types per path in alphabetical order buries it among the normal ones. Show each type with its share of traffic and the date it was first seen, and the reviewer spends the time where the risk is.

On the wildcard, you can keep part of the signal without enumerating keys. Track the value types seen under it as one set, plus a daily count of distinct keys. A partner that starts sending a nested object where it used to send strings still alerts, and a key count that jumps from 40 to 40,000 tells you someone has started putting request IDs in metadata.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Rarest first is right and alphabetical would have buried the one entry the review exists for. Taking the traffic share and the first-seen date with it.

The hole rarest-first opens is the long tail, because the rarest entries are mostly not interesting. One corrupt payload from a partner's bad deploy. A truncated field from a timeout. A test record somebody sent at a live endpoint. If the reviewer works upward from the rarest, the acquirer sending amount as a string sits at position four hundred behind three hundred and ninety-nine one-offs, and the review dies of boredom before it arrives.

The split that fixes it is not frequency, it is recurrence. A type at 0.1 percent that appears on most days of the sample is a partner's normal behaviour and needs a decision. A type at 0.1 percent that appeared twice on one afternoon is an incident and needs a different person. Identical rarity, opposite meaning, and days-seen separates them without anyone exercising judgement.

On the wildcard I would take your two signals and change one thing. Alert on the shape of the growth rather than the level, because distinct key counts grow legitimately and any fixed threshold is noisy at forty and useless at forty thousand.

And I would put a hard cap at the edge alongside the alert. Forty thousand distinct keys under a metadata object is not only a monitoring problem. It is an index somebody will try to search in eighteen months, and it is a place personal data ends up precisely because nobody declared a schema for it. Rejecting it at the boundary is the cheapest that decision will ever be. It is also the only point in the system where we can still say no without breaking a partner who has come to depend on it.

Collapse
 
mickyarun profile image
arun rajkumar •

@anp2network @peterbuildssecure @salparvez @_firelinks — I have written this thread up, and you are all quoted at length.

The part I could not stop thinking about: every move in the twenty-reply branch deletes a store and creates one somewhere else, and at the bottom there is a retention number that nobody in the thread could derive. Not for lack of trying. Payments does not derive it either. It gets handed one by a regulator, and the reason that works has nothing to do with the number being right.

dev.to/mickyarun/two-strangers-bui...

@anp2network, your pushback about a public revision history buying most of what a regulator buys is in there, along with where I think it stops. @peterbuildssecure, the two-tier retention split is the spine of the middle section. @salparvez, I used your roof-and-foundation version because it is the one I actually remember. @_firelinks, the circular-canary point is in the credits and it deserved its own article.

Corrections welcome, as usual. You have a decent record on that.

Collapse
 
road511 profile image
Roman Kotenko •

The line I'd add to yours: a guard that never fired, a guard that silently stopped running, and a guard that is running perfectly while watching the wrong property all produce the same green.

We poll about sixty public road-traffic feeds. A German state feed served the same bytes for thirteen hours under 200 OK. Every fetch succeeded, the error counter stayed at zero, and the row count stayed exactly right — because our guard was a COUNT guard. A feed that keeps returning the correct number of stale things is invisible to it by construction. Nothing was broken, nothing was skipped, nothing was misconfigured. The check ran on schedule and answered the question it was built to answer, which was the wrong question.

Your two cases are both "is it running". This one is "what does it actually read", and no amount of verifying the guardrail is alive would have surfaced it. The thing that did surface it was boring: writing down, per guard, the sentence "this will not catch ___". The COUNT guard's sentence turned out to be "anything that changes the shape of the data rather than its volume" — which is most of what goes wrong with a third-party feed.

The part I didn't expect is that the fix has a blind spot you have to keep. The replacement compares the newest timestamp inside the payload against wall clock. That correctly fires on a frozen feed, and it also fires on snowplough and seasonal-closure feeds, which stop moving every summer and are supposed to. There is no version of that check that is both complete and quiet, so it ships with an explicit exemption list — and the exemption list is the honest part of it, not the embarrassing part.

(Disclosure: I build a commercial road-data API, so feeds that lie politely are the day job. Not pitching anything — the "never fired vs stopped firing" distinction is what I'm taking away.)

Collapse
 
mickyarun profile image
arun rajkumar •

The third case is the one my article doesn't cover, and you're right that no amount of liveness checking reaches it. A guard watching the wrong property is alive, on schedule, and green.

The payments version is a settlement file that arrives on time with the right row count and yesterday's contents. Every check passes. Counts reconcile. What catches it is never the pipeline — it's someone downstream noticing a number that should have moved and didn't.

Your "this will not catch ___" sentence is the best thing in this thread. It does the work a threat model is supposed to do and almost never does, because it's scoped to one guard and it's one line. I'm stealing it.

The exemption list being the honest part is where I'd push slightly. It's honest the day you write it. Snowplough feeds are a stable exemption. The risk is the list growing by one every time the check is noisy, and six months later nobody can say which entries were reasoned and which were added to stop a 3am page. Does yours carry a reason per entry, or is it a set of feed ids?

Collapse
 
road511 profile image
Roman Kotenko •

Per entry, enforced — but your prediction is already half true in our list, so here is the actual state rather than the design intent.

Context first, because it decides how much of this generalises: the thing being guarded is an aggregation layer over government road-traffic feeds — 896 polled endpoints across 285 upstream servers on two continents, every one of them somebody else's publishing decision that can change shape without telling us. The normalised output is what somebody pays for, so a feed that freezes is not an internal annoyance; it is a customer being served yesterday's road closures and having no way to tell.

On the question. It is not a set of ids. It is two columns on the resource row: an expiry timestamp and a reason, with a CHECK constraint that rejects the timestamp unless the reason is non-empty. So "add the feed id to shut it up" is not reachable; the schema makes you type something.

The expiry is the part I would argue for harder than the reason. An exemption is a blindfold, and a blindfold with no end date is how one of our weather-station feeds stayed dead for roughly eight months — silenced once, never reviewed, and nothing in the system had a reason to look at it again. Now the timestamp passes and the feed starts alerting on its own, with no cleanup step anybody has to remember.

Now the part that proves your point. I pulled the live list before writing this: 10 exemptions, all still in date, none missing a reason. But 7 of the 10 carry the same 153-character string, written in one batch sweep five days ago — "verified empty upstream, peers of the same type still producing". Only 3 have a reason specific to that feed, and those are the seasonal ones where somebody actually went and looked (a plough feed whose timestamp froze in April; a winter-roads endpoint serving 767 segments that all read "No Active Reporting").

So the constraint buys the weaker half. It can force a reason to exist; it cannot force the reason to be about that entry. A batch sweep satisfies it perfectly and produces exactly the list you describe — one where nobody can later tell which entries were reasoned. The expiry is what saves it, and only because it is short: those 7 all fall due on 13 October, and re-typing the same sentence seven times is annoying enough that somebody will either look properly or delete them.

Which suggests the rule is not "carry a reason" but "carry a reason and a date, and keep the date short enough that renewing is more expensive than checking". If renewal is cheap, the reason rots and the date does nothing.

Your settlement-file example is the cleanest version of this I have seen — right row count, right arrival time, yesterday's contents. Same shape as ours: the payload is well-formed and on schedule, and the only thing wrong with it is that it is the previous one. Nothing in the transport layer can see that, which is why it always gets caught downstream by a human noticing a number that should have moved.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Pulling the live list before answering is worth more than the design, and 7 of 10 sharing a 153-character string is the number I'll remember from this thread.

The rule you've ended on is sharper than mine and I want to state it back: a reason and a date, with the date short enough that renewing costs more than checking. The constraint can only make a reason exist. Pricing is what makes it true.

13 October is the experiment. If those seven come back with the same sentence, renewal was cheaper than looking and the date did nothing. One cheap tweak before then: reject a renewal whose reason matches the previous one for that row. It doesn't force honesty, but it forces a second keystroke, and the batch sweep stops being a paste.

The eight-month weather-station feed is the case that makes the expiry non-negotiable. An exemption with no end date is a guard you deleted without telling anyone. Same green.

I'd like to write this up properly, with the 896 endpoints and the 7-of-10 finding, credited to you. Shout if you'd rather I didn't.

Thread Thread
 
road511 profile image
Roman Kotenko •

Yes — write it up, and thank you for asking rather than just doing it.

Two things to take with you, both re-pulled today (2026-09-21) so the article starts from live numbers rather than the ones in my comment three days ago.

897 polled endpoints across 286 upstream servers — NA 674 / EU 223, and "server" there means one with at least one enabled resource, which is the definition that makes the number reproducible. On 18 September those were 896 and 285. That drift is the point: ±1 over three days, so a figure with a date on it stays true for a useful while, and a figure without one quietly stops being true. Whatever you use, please stamp it.

One correction to what I told you, because it will read wrong otherwise. The 10 exemptions are the North American deployment. The EU one has zero — not a better-disciplined team, just a younger list. I said "the live list" and meant one of two, which is exactly the kind of thing that survives into somebody else's article as a fact about the whole product.

The 7-of-10 split is unchanged: same 153-character string, same 13 October expiry, three feed-specific reasons due 15 November. Length is not the tell, incidentally — the three honest ones run 199, 210 and 154 characters. The only thing separating them is that they are about their own feed.

Your renewal-dedup idea is the sharper half of this, and I want to be careful not to repay it with a promise. Rejecting a renewal whose reason matches the previous one for that row cannot force honesty, as you say — but it does convert a batch sweep back into seven separate decisions, and seven is where somebody gives up and looks. I am not telling you it ships; I am telling you it is the cheapest idea anyone has put on this list, and that it is now written down somewhere it will be read.

And on 13 October: I will try to come back here with what happened to those seven, whichever way it goes. Not a promise — the honest version of a commitment to report is that people forget, and a date in a comment thread has no CHECK constraint behind it. But we are the ones holding the data, so if the result gets published at all it should be by us, and it should be the result rather than the intention.

Thread Thread
 
Sloan, the sloth mascot
Comment deleted
 
road511 profile image
Roman Kotenko •

That's all of it correct — nothing further from me. If you want the 21 September figures re-pulled on the morning you publish, say the word; the drift is small, and it is the whole argument for stamping them.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Yes — re-pull on the morning, and I will take whatever the numbers are then rather than the ones I have.

One correction on my side: the long reply you answered went out from the wrong profile. Mine, wrong account, my mistake. Flagging it because you are crediting a handle in the write-up and it should be this one.

Collapse
 
reidmarlow profile image
Reid Marlow •

The negative control canary is the only pattern that reliably catches parser drift in policy filters. When an agent runner updates its tool call schema or changes how whitespace is stripped from shell arguments, argument-matching regexes often stop matching the payload structure entirely. Because the parser sees no blacklist hits, every command passes through.

I ran into this after a runtime update changed JSON serialization on bash arguments. The filter silently evaluated every destructive command as benign because the regex was looking for a string pattern that no longer appeared in the raw input. The CI suite stayed green because nothing failed explicitly.

The fix that held up was bundling a synthetic poison payload into every filter test run. If the gate fails to reject the known bad command, the test harness hard-errors immediately. Treating a guardrail that never rejects anything as a test failure stops parser rot before code reaches production.

Collapse
 
mickyarun profile image
arun rajkumar •

The part that makes this hard is that the poison payload is written in the same format the parser stopped understanding. If a runtime change moves bash arguments from a string to a structured object, a canary authored against the old shape goes green-because-unmatched in exactly the way production did. It fails to reject, but so does everything else, and from inside the test those two look the same.

So the canary shouldn't be a literal. It should be constructed by the same serialiser the real call path uses. Then a format change breaks the canary's construction, loudly, instead of quietly changing what it means.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

The negative-control framing is the right import from lab science. The other half we found necessary: a result without provenance isn't a result. If an eval says pass but can't tell you which inputs it ran, when, and against which version, it is indistinguishable from a grader that stopped rejecting. We handle it by treating every result as a claim with a source and a verification state, and deriving a review queue over everything that hasn't fired recently — so the guardrail that never fires shows up in the queue instead of disappearing into green.

Collapse
 
mickyarun profile image
arun rajkumar •

Provenance is what turns a result into something you can argue with. The dimension I'd add to the claim is which side of the boundary moved. We read from bank APIs where the schema changes on the provider's release calendar and not on ours, so "this check passed" and "this check passed against the contract we last read" are different sentences, and only the second one survives a quiet provider change.

A review queue over checks that haven't fired recently is the right shape. I'd sort it by how fast the thing being checked is changing underneath, not by how long the check has been silent.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

"Which side of the boundary moved" is a state we had to name separately. A stamp on a check is bound to the content it verified, a hash of the contract as last read, so when the provider ships a new schema the stamp lapses on its own, without anyone re-running anything. Lapsed and unverified are different rows: lapsed means the ground under a verified result moved, unverified means it was never verified.

The queue is derived in that order, quarantined › lapsed › unverified › awaiting-stamp, which is your sort. Rate of change underneath outranks duration of silence, because a check that has been quiet for a year against a contract that hasn't changed is still evidence, and a check that passed yesterday against a contract that changed this morning is not.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Binding the stamp to a hash of the contract as read is the part I'd steal. What I'd be careful about is what goes into the hash. Our banks ship schema revisions on their own calendar and most of what moves is fields we never touch, so hashing the whole document lapses everything at once and the queue becomes a chore people clear without reading. Hash the projection you actually depend on and a lapse means something.

Quarantined before lapsed before unverified is the right order. It's also the first time I've seen someone put awaiting-stamp at the bottom instead of treating it as an error.

Thread Thread
 
salparvez profile image
Sal Parvez | ML Systems •

Hash the projection, not the document: taking that. A lapse that fires on fields nobody depends on trains people to clear the queue without reading it, which is the guardrail failing in a new costume. Ours binds the stamp to the content of the entry, not the whole record, for the same reason; a re-shingled roof should not lapse the foundation. And yes, awaiting-stamp at the bottom is on purpose. It is the normal state of a new claim, not an error, and treating it as one is how queues turn into noise.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Re-shingled roof and the foundation is the better version of what I was groping at.

The case that still gets past both of us: a field that does not change shape and does change meaning. Same name, same type, same position in the projection, and the bank quietly starts populating it from a different source. The hash is identical. The stamp stays valid. Nothing lapses, and the thing you depend on is now wrong.

I have no mechanical answer for that one. What we do is cheap and manual. The projection carries a note on where each field is supposed to come from, and a human reads it when a bank ships a release note. It catches the ones that get announced. It catches none of the ones that do not.

Collapse
 
utilvance profile image
Utilvance •

The review queue over recently-silent guardrails is the part most teams skip. They build the check, they build the dashboard, and they stop there. Surfacing the ones that have not fired recently as a separate queue flips the default from "trust until broken" to "verify until confirmed." That is a much safer baseline, especially when the guardrail count grows and no single person has the full picture of what each one is supposed to catch.

Collapse
 
mickyarun profile image
arun rajkumar •

"Verify until confirmed" is the right default and the queue is the right shape. Where it gets hard is the word "recently".

A guardrail that has not fired in thirty days is either dead or doing its job in a quiet month. The queue cannot tell those apart, so it lists both, and the ratio decides whether anyone reads it. Most guardrails on most systems are legitimately silent most of the time. So the queue fills with things that are fine, somebody clears it in bulk, and six weeks later the queue is the check nobody reads.

Salparvez's negative control is what fixes that, and I think it has to be paired with the queue rather than offered as an alternative to it. Fire the guardrail deliberately on a schedule, with an input you know is bad. Now silence means something specific: the synthetic did not fire either. The queue stops being "has not fired lately" and becomes "did not fire when we made it".

The cost is that every guardrail then needs a known-bad input somebody maintains, which is real work, and is why most teams stop at the dashboard.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Agreed, and the sort order matters as much as the queue. Silence alone ranks a check that has not fired because nothing changed the same as one that has not fired because it is broken. Arun's point above is the fix: sort by how fast the thing under the check is moving. Then "verify until confirmed" becomes a schedule instead of a slogan, and the queue stays short enough that someone actually reads it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Sorting by how fast the thing under the check is moving is the part worth keeping. It also gives the queue a natural size, which matters more than it sounds. A review queue that never empties gets ignored in about three weeks.

One refinement: rate of change of the code under the check, not of the traffic through it. Flat traffic is often the reason nothing fired, so traffic-based ranking buries exactly the checks you want to look at. A deploy that touches the path is the thing that should push a silent check to the top, and that signal costs nothing. You already have the commit.

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

The dead man's switch framing is exactly right, and I'd push it one step further: the failure mode isn't just "nobody checks if the guardrail ran," it's that most teams can't even answer what "running" means for a check that's supposed to be silent 99% of the time.

We hit this building webhook delivery infra at MSG91. A retry policy that never retries looks identical to a retry policy handling everything perfectly, until the one week your provider's API silently starts 200-ing with empty bodies and nothing downstream ever complains because nothing downstream expected a failure to look normal.

The fix that actually worked for us wasn't a better guardrail. It was making the guardrail's silence itself an anomaly. If your negative-control canary hasn't fired in N days, that absence pages someone, same severity as an outage. Cheap to build, almost nobody does it, because it feels like alerting on nothing happening.

Your point about agents shifting the ratio is the part people will underrate. The whole pitch of agentic review is "more automated checks, fewer human eyes." Nobody's pricing in that the automated checks now need their own review layer, and that layer usually doesn't exist until after the first quiet failure gets expensive.

Collapse
 
mickyarun profile image
arun rajkumar •

The 200 with an empty body is the exact one. There's a payments version of it where a status callback arrives, parses cleanly, and carries nothing that lets you tell success from silence. The retry logic has no reason to fire, so it doesn't, and the graph stays clean.

Alerting on absence feels wrong to build and is the only thing that works. The usual objection is noise, but it's only noisy if you pick N by gut. Set it from the observed inter-arrival time of the last hundred rejections and it stays quiet until the distribution actually moves.

Your last paragraph is the part I'd want people to take away. More automated checks means more unwatched checks, and nobody budgets for the watching layer until the first quiet failure gets expensive.

Collapse
 
tejas_shinkar profile image
Tejas Shinkar •

The "score with no provenance" section nails something I hadn't put words to. A pass with no explanation feels identical to a real pass until the one time it actually mattered and nobody could say why. Curious if you've seen teams solve this without just bolting on more logging , feels like the kind of thing that needs to be designed in from day one, not patched after the fact.

Collapse
 
mickyarun profile image
arun rajkumar •

Not really, no. The one place we got it right was designed in, and only because money forced it. Everything else is bolted on.

But the distinction I'd draw isn't logging versus not logging. It's whether the result carries the version of the rule that produced it. That's one field. Cheap to add on day one and impossible to backfill, because the old rule versions are gone. So the thing to design in isn't a logging system, it's the field.

Collapse
 
to21as profile image
Tobias •

The zero-cases one got me almost exactly as you describe it. Three scheduled collector runs in a row wrote zero rows and every dashboard stayed green, because the heartbeat fires when the run finishes, not when it collects anything. The run was alive. It just had nothing to say and no way to say so.

What I added afterwards is your run-count point as a hard failure rather than a metric: a row count below one is an exit code, not a line in a log nobody reads.

Have you got the last-rejection date running somewhere in practice, or is it still the thing you would like to have? That is the one I would expect to quietly stop being updated.

Collapse
 
mickyarun profile image
arun rajkumar •

Honest answer, partly. On the payment side we alert on absence, because a callback stream going quiet is indistinguishable from every payment succeeding, and that one has a cost attached, so it got built. The last-rejection date on lint rules and policy checks is the thing I'd like and don't have. Nobody funds a dashboard column for a check that has never been wrong.

Your row-count-below-one as an exit code rather than a metric is the version I can actually ship, because it needs no new system. Just a stricter definition of what finishing means.

Collapse
 
routinekit profile image
RoutineKit •

This matches the freelance version I keep hitting: the “process” exists in a Notion page nobody opens on the day it matters.

What stuck for me was turning the guardrail into the last step of the work itself — same chair, same slot — not a separate audit. If kill/re-steer isn’t how the task ends, it becomes optional and optional dies under deadline pressure.

Do you treat the check as a calendar ritual, or as a hard stop baked into done?

Collapse
 
mickyarun profile image
arun rajkumar •

Hard stop where money moves, ritual everywhere else, and the honest half is that the ritual side decays exactly the way you describe. Optional dies under deadline pressure is the article in one line.

What made the difference wasn't discipline, it was that the check sat inside the only path the operation could take, so skipping it meant not doing the work at all. A Notion page can't do that. Your four-line header is the same idea moved earlier in the process, and I left a longer note on that post.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.