DEV Community

Cover image for You restart the server and the agent loses all ten tasks
Guillermo Leyendeker
Guillermo Leyendeker

Posted on Originally published at leyendeker.com

You restart the server and the agent loses all ten tasks

You leave the agent working through a queue of ten tasks. You need to deploy a server improvement, you restart it, and all ten are gone.

Not just the pending work: also any trace of which one was running and what state it was left in. The queue lived in the process's memory, so any restart — a deploy, an update, a power cut — wiped it entirely.

And the server you're restarting is the very one you're improving every day. Every improvement forced a choice: wait for the queue to drain, or lose the plan.

The fix isn't "put the queue in a database" and done. What you actually have to decide is what a half-finished task means, and who gets to change its state when the process comes back up.

How it used to work

The original queue was an in-memory list. You ticked the tasks you wanted to run on screen, the order was the order you'd selected them in, and the system processed them one by one.

It worked as long as nothing interrupted it. With two serious problems.

The first is the one above: every time I wanted to deploy an improvement to the orchestrator, I had to choose between waiting for the queue to drain or losing the plan.

The second, quieter one: the first failure aborted the whole queue. An error in the third of ten tasks cancelled the remaining seven. And those seven had nothing to do with the one that failed: they were simply behind it in line.

The redesign

The queue moved to the database, with one design decision I particularly like: position is a real number, not an integer.

That allows inserting between two elements without renumbering the whole list. If you want to slot something between positions 3 and 4, you assign it 3.5. It's an old technique and it solves at the root the problem of reordering a persisted list without writing fifty rows every time someone drags one.

The order is curated by hand, mixing fixes, features and audits into whatever sequence makes sense that day. You reorder by dragging, as long as tasks are still waiting. Whatever is running is pinned and locked, which is the bare minimum: you can't reorder something that already started.

What to do when a task fails

There was a product decision disguised as a technical detail here, and I resolved it by not resolving it: I made it a per-run option.

You can choose for the queue to carry on with the rest, flagging the failed one, or to stop and leave everything else waiting. Both positions are defensible depending on the case. If you're processing ten independent fixes, you want it to continue. If the third failed because the environment broke, you want it to stop before the remaining seven fail for the same reason.

Later I added a stricter rule that turned out to be the most useful in practice: the queue doesn't resume past an unresolved failure. You can continue, but someone has to have looked at that failure and decided what to do about it. Without that, the queue stays put.

It's deliberately inconvenient. The alternative — always continue — lets failures pile up unseen, which is exactly the problem behind the 1,771 unprocessed incidents I described a few entries back.

A single execution mechanism

A side effect of the redesign I hadn't anticipated: the "run everything" and "repair everything" buttons stopped being separate mechanisms.

Before, each was its own pipeline, with its own advancement logic and its own bugs. Now they're simply shortcuts that bulk-enqueue with a preset order and start the queue. One execution mechanism, all of it visible and reorderable.

It's the kind of simplification that appears on its own once the right abstraction finally exists. While the queue was ephemeral, having parallel pipelines seemed reasonable. With a persistent, curatable queue, it makes no sense at all.

Actually surviving a restart

Persisting the queue solves half the problem. The other half is what happens to whatever was running at the moment of the restart.

That task lands in limbo: the database says it's in progress, but the process executing it no longer exists. If nobody does anything, it sits there forever, occupying the head of the queue and blocking everything behind it.

In August I added the two missing pieces: graceful shutdown, so in-flight runs are marked before dying, and startup reconciliation, which checks what's declared as running with no process behind it and resolves it.

And a rule that seems obvious and took me a while to find: re-read the task's status before launching its row. Because between enqueuing and its turn arriving, anything could have happened — the task might have been completed by another path, or postponed. Launching blind against a stale snapshot is a constant source of duplicated work.

Why this sums up the whole project

Of everything I built over these months, the persistent queue is what best captures the thesis: in a system with agents, the state of in-flight work matters as much as the work itself.

An in-memory queue works perfectly until the first restart, and the first restart always arrives at the worst moment. It isn't an optimisation: it's the difference between a tool you use while you're watching and a system that runs when you're not.

Today the queue is the system's only execution path. There's no way to run a task outside it, and that — which felt like an annoying restriction at first — turned out to be what makes everything else observable.

With the queue surviving restarts, the system could run on its own for hours with me nowhere near it. And that raised a question I'd never asked, because I'd never let it run that long: how long to let it think before interrupting. That's what the next entry is about.

Top comments (0)