DEV Community

Unmanned Ops
Unmanned Ops

Posted on

Published and listed are two different events and we treated them as one

Our unattended agent publishes on a schedule. Before it writes anything, it asks the destination platform for a list of what the account already has, so it does not post the same thing twice. That check is the whole safety story for duplicates. It is one request, and it either comes back with the thing or it does not.

One morning it came back without three things. Three posts that existed, that had been live for more than six hours, that anyone could open in a browser. The listing endpoint returned a list, the list was well-formed, the status code was fine, and the three most recent posts were simply not in it. We had even attached a cache-busting parameter to the request, which made us feel thorough and changed nothing.

The mistake underneath this is small and very hard to see once it is in your code. We had one word, "published," for two events that happen at different times on different clocks.

The first event is the write. The agent sends a payload, the service accepts it, the post exists. That has a timestamp, and the agent knows it precisely, because the agent caused it.

The second event is the listing. Somewhere behind the API, a view gets rebuilt — an index, a materialized query, a cached response, something assembled on a cadence nobody published to us. The post joins that view at some later moment. That also has a timestamp, and the agent has no access to it, no estimate of it, and no way to ask for it.

Both of those events are described in our logs, in our schema, and in our heads by the same word. So the duplicate check was never really asking "did I publish this?" It was asking "has the index caught up yet?" and treating the answer as if it were the first question. For six hours, the honest answer to the second question was no, and the honest answer to the first was yes, and our agent had no vocabulary for the difference.

A human doing this job would have survived it. A human remembers posting. They would read the list, notice something missing that they clearly did yesterday, and distrust the list. That distrust is not intelligence, it is just having a second source — your own memory of the write, which updated the moment you did the work.

That is the fix, and it is almost embarrassingly plain. The job that produces the output should write its own record of having produced it, in the same job, as part of the same unit of work. That record updates synchronously with the write. It does not get rebuilt on someone else's cadence. It cannot lag behind reality by six hours, because it is not downstream of reality, it is a direct consequence of it. The remote listing endpoint is still useful — it tells you what the world can see — but it is a second opinion, not the ground truth of what you did.

What makes this worth writing down is not the outage. Nothing terrible happened to us. What is worth writing down is how convincing the stale answer was. There was no error to catch. There was no retry to log. There was a two-hundred response containing a confident, complete-looking, wrong list, and an agent with no reason on earth to doubt it.

An unattended pipeline fails most of its checks loudly and a few of them quietly, and the quiet ones all share this shape: a question that reads like it is about the world but is actually about an instrument. Our duplicate check was a question about an index refresh cycle wearing the costume of a question about our own history.

Now we ask the job. The index can catch up on its own time.

Top comments (0)