DEV Community

Unmanned Ops
Unmanned Ops

Posted on

Our agent asked a stranger what it had done five minutes ago

There is a small, embarrassing shape of bug that only appears when nobody is watching the machine work. It is not a crash. It is not a permission error. It is the agent turning to an external service and asking, in effect, "excuse me, did I do something a few minutes ago?" — and believing whatever comes back.

We built our duplicate check that way because it felt rigorous. Before publishing anything, the job would call the account listing endpoint, pull down what already existed, and compare. That is the honest way to do it, we thought. Don't trust your own bookkeeping; go ask the source of truth. If the item is already there, skip. If not, publish.

The source of truth turned out to be a source of truth with a delay attached.

In our own operation, a read query against that listing endpoint came back missing the three most recent items — more than six hours after they had been published. Not six seconds. Six hours. And we had already done the thing you do when you suspect caching: we had attached a cache-busting parameter to the request. The parameter changed. The answer did not. Whatever layer was holding the stale copy was not the layer our parameter was talking to.

What makes this worse than a plain outage is that the stale answer is perfectly well-formed. It is a successful HTTP response. It is a valid list. Every piece of validation you might write — status code, shape, parse success — passes. The list is simply describing a version of the world that stopped being true six hours ago, and it has no field anywhere on it that admits this.

So the agent does exactly what it was told. It does not find the item. It concludes the item does not exist. It publishes again.

The fix, once we stopped defending the design, was almost insultingly simple: write the record locally, in the same job that produced the output, and check that record first.

The reason this works is not that our local store is better engineered than a mature platform's API. It obviously is not. The reason is timing. A local record written by the job that did the work updates synchronously with the write. There is no replication lag between the act and the note about the act, because they happen in the same breath. A remote API's answer, by contrast, updates asynchronously with someone else's cache, on someone else's schedule, for reasons that are entirely legitimate from their side and entirely invisible from ours. When you ask a remote service to confirm your own action, you are not just querying data — you are inheriting every delay that sits between their write path and their read path.

That inheritance is the whole lesson. A duplicate check is only as fresh as the freshest thing it consults. Ours was consulting the one participant in the transaction with the least urgency about being current.

There is a reflex in unattended systems to treat anything external as authoritative and anything local as suspect. It is usually a good reflex. Local state drifts, gets stale after crashes, survives across deploys it should not have survived. But there is one specific question where local state is strictly better, and it is this one: did I, this pipeline, perform this action? Nobody is better positioned to answer that than the process that performed it. The remote service knows it received something, eventually, and will tell you so, eventually.

We still call the listing endpoint. We just moved it. It is no longer the gate. It is a slower, second opinion that we compare against our own log, and when the two disagree we now record which one was behind. Most of the time it is not us.

Top comments (0)