DEV Community

Cover image for The harness is one integer column
Raj Murugan
Raj Murugan

Posted on Originally published at rajmurugan.com

The harness is one integer column

A background job was re-billing an AI model every sixty seconds, forever. Not a rewrite. Not a redesign. The fix was one INT column, capped at five. And the fix itself shipped with a gap: it forgot to log why a row gave up, so the very rows it saved became impossible to explain. This is what I now do differently, after both halves of that day.

Diagram of a bounded retry: a poison-pill row polled forever re-bills a model on every lap until an atomic database counter, incremented before each risky attempt and capped at five, flips the row to abandoned and logs why it gave up.

The failure, in one breath

A row that can never reach its terminal state, sitting behind a poller that checks it every tick, is an infinite retry loop wearing a queue's clothes. Nothing was broken in the traditional sense: no exception, no crash, no alarm. A background close job kept picking up the same "in progress" row, calling the model again, failing to close it again, and handing it back to the next poll. Every lap re-billed the model. Nobody was watching for a row that could not finish, because "could not finish" is not an error, it is just still running.

This is a level-triggered trigger against a condition that never clears. An edge-triggered retry fires once on the transition into failure. A level-triggered poller fires every time it observes the bad state, and if nothing ever changes that state, it fires forever. The row was the poison pill. The poll was the mouth that kept swallowing it.

Why the counter belongs in the database, not memory or metrics

The instinctive fix is a counter: retry, but only a few times. Where you put that counter decides whether it works.

Not in Lambda memory. A retry counter held in the function's own memory resets on every cold start, and a poison-pill row cold-starts its handler just as often as anything else. The counter never reaches five, because it is reintroduced to zero constantly.

Not in a CloudWatch metric. Metrics observe. They tell you a number happened. They do not gate the next attempt, and a dashboard with a suspicious spike on it does not stop the next poll from firing. Watching is not the same instrument as stopping.

What actually holds a ceiling under concurrency is a durable, atomic transition in the database itself, and it has to do two jobs, not one: stop the count from being lost, and stop two pollers from both grabbing the same row on the same tick.

-- Claim the row and count the attempt in one statement. Zero rows back means
-- someone else already has it, or it is already at the cap. Either way, skip
-- it this tick.
UPDATE jobs
SET retry_count = retry_count + 1,
    state = 'claimed'
WHERE id = $1
  AND state = 'pending'
  AND retry_count < 5
RETURNING retry_count;
Enter fullscreen mode Exit fullscreen mode

The state = 'pending' clause is what stops two poller instances from both picking up the same row and both calling the model, not the increment on its own: an increment alone only protects the count, it does not claim the row. The retry_count < 5 clause is what makes the cap a fact the database enforces, rather than something application code has to remember to check after reading the count back. (This is the behaviour under the default Read Committed isolation Postgres and Aurora ship with; under a stricter isolation level a second concurrent writer gets a serialization error instead of quietly waiting its turn, and needs its own retry to match.)

If the risky work then fails, a second atomic statement decides what happens to the row next:

UPDATE jobs
SET state = CASE WHEN retry_count >= 5 THEN 'abandoned' ELSE 'pending' END,
    last_close_error = $2
WHERE id = $1
RETURNING state;
Enter fullscreen mode Exit fullscreen mode

Under the cap, the row goes back to pending for the next poll. At the cap, the same statement flips it straight to abandoned, and the poller's own state = 'pending' clause means it will never claim that row again.

The bounded-retry pattern

Three decisions make this actually work, and each one is a place the naive version gets it wrong.

Claim and increment in the same statement, before the risky work, not after. If you increment on success or on a clean failure, a crash mid-attempt, which is exactly what a poison pill causes, never gets counted, and the cap never bites. Count attempts, not successes.

Force a terminal state at the cap, enforced by the database, not read back and checked by application code. At five, the same WHERE retry_count < 5 that gates every claim stops matching, and the follow-up statement flips the row to abandoned. The retry loop has an exit now, one that does not depend on some later code path remembering to check.

Do the arithmetic on the worst case, honestly. Five attempts at the model's per-call cost is a bounded, small number, call it a dollar. Uncapped, the same bug run for a year at one call a minute is the same unbounded shape as any retry loop nobody put a ceiling on: a small per-call cost, multiplied by nothing ever stopping it. The whole value of the cap is that it turns an open-ended bill into a number you can write down in advance, whatever that number turns out to be for your own per-call cost.

The fix needed its own observability

Here is the part I would tell you to do differently if I were doing it again. Version one had the claim, the cap, and the terminal state. It did not have the last_close_error column shown above: it flipped rows to abandoned at five and wrote nothing about why.

So the abandoned rows carried a null reason. They were safe, in the sense that they had stopped costing money. They were also undiagnosable: nobody could look at a batch of abandoned rows and tell you whether they were all the same bug, five different bugs, or a symptom of something upstream getting worse. Was recovering them safe? Nobody could say, because nobody knew what had actually gone wrong on attempt one through five. last_close_error was a follow-up migration, added after that gap got noticed, not part of the design from day one.

A guardrail with no telemetry is a black box you will have to re-debug from scratch the next time it fires, using none of the information the guardrail itself was sitting on the whole time.

The same shape, one layer up

None of this is specific to a background job. Swap poller for agent loop and the mapping is exact: bound the iterations, force a terminal state at the bound, and log why it stopped. I have written before about what happens when an agent loop skips all three: a runaway that ran for days because nothing capped it, and would have been just as undiagnosable as this one if it had been capped without the third piece.

The stopping half is the part everyone remembers to build. The observability half is the part that gets cut when the deadline is close, because a capped, silent failure still looks like the incident is over. It is over in the sense that the bill stopped. It is not over in the sense that you can tell anyone what actually happened, and the next stuck row, or the next runaway agent, gets the exact same treatment: caught, capped, and still a mystery.

What I now do

Four rows: the counter lives in the row itself, not Lambda memory or a CloudWatch metric; increment before the risky work so the cap counts attempts; force a terminal state at the cap so the loop has an exit; log the reason on every attempt, not just the last one.

A retry ceiling lives in the row it protects, incremented atomically, not in a process's memory and not only in a metric.

Claim and count in the same statement, gated by the cap, so the database enforces the ceiling instead of application code checking it after the fact.

A cap without a logged reason is half a guardrail. Write down why it gave up on every attempt, not just the last one.

Do the worst-case arithmetic before you ship the cap, honestly, so you know the number you are actually bounding.

If a bounded retry in your system does not write down why it stopped, what would it cost you to find out, the next time it fires?

This pairs with Every dashboard was green while the agent burned six figures a year, the incident that made me go looking for this pattern everywhere else. Find me on LinkedIn or via rajmurugan.com.

Top comments (0)