DEV Community

Remdore
Remdore

Posted on AI-assisted

Pausing an agent mid-task and resuming it four minutes later, with its memory intact

Most of the agent-hosting products that turned up this year solve the same first problem, which is that nobody wants a model running rm -rf against their laptop, so the model gets a container somewhere else instead. That part has become commodity. The part that has not, and the part I wanted to check properly, is what happens to a long-running agent when you stop paying attention to it halfway through its work.

DigitalOcean's Managed Agents went into public preview recently, and the documentation makes a claim that is stronger than it first looks. Pausing a session, it says, preserves the processes, the memory and the workspace filesystem, and resuming brings them back. A stopped container loses everything that was not written to a volume, so if that sentence is literally true it is a different kind of thing, and I could not find anybody who had gone and tested it. So I spent a morning and about seven cents finding out.

Designing a test the filesystem cannot fake

The obvious version of this test is worthless. If you write a file, pause, resume, and read the file back, you have proven that a disk survived, which was never in doubt. The claim about memory needs something that lives only in memory and is never read back from anywhere.

So I wrote the dumbest possible process. A bash loop holding a counter in a shell variable, incrementing once a second, appending the current value and a timestamp to a log. The log is write-only from the process's point of view. Nothing ever reads it back, so if the session were destroyed and recreated with the filesystem restored, the counter would start again from one and the old log would simply have a new sequence appended to it.

#!/bin/bash
i=0
while true; do
  i=$((i+1))
  echo "$i $(date -u +%H:%M:%S)" >> /tmp/tick.log
  sleep 1
done
Enter fullscreen mode Exit fullscreen mode

Starting it needed a bit of care, because the exec channel into the sandbox kills its children when it closes. setsid nohup /tmp/tick.sh </dev/null >/dev/null 2>&1 & disown was what survived.

Then I let it run for about a minute, paused the session, went and made coffee, and resumed.

What came back

Two consecutive lines in the log, which is the whole result:

47 08:05:42
48 08:10:10
Enter fullscreen mode Exit fullscreen mode

Forty-seven, then forty-eight, on the same process at PID 590 holding the same shell variable it had before, and the only evidence that anything happened at all is the four minute and twenty-eight second hole where a one-second tick should have been.

The in-memory counter across the pause

The pause call itself returned in 0.86 seconds and the resume in 1.16. Neither of those numbers is doing much work, since the interesting quantity is the four and a half minutes in between, during which the session was not consuming anything.

I ran the same check against the agent's own context rather than a shell variable. Before pausing, I had the agent generate a random ticket identifier, HARBOUR-7742, and write it into a file. After the resume I asked it what the ticket was called, and it answered from its conversation history in eight output tokens without touching the filesystem. Both kinds of state came back, the operating system's and the agent's.

The sandbox is a real machine

Worth confirming what the process was actually running on, since "sandbox" covers everything from a chroot to a VM.

Hypervisor detected: KVM
CPU: 2 vCPU   Memory: 3939 MB   Disk: /dev/vda   Kernel: 6.1.176
Enter fullscreen mode Exit fullscreen mode

A microVM with its own kernel, its own block device and hardware virtualisation underneath it, not a namespace on a shared host. Session creation from the API call to status READY took 15.97 seconds, which is slower than a container and about what a Firecracker-class VM costs you. Sizes run from mars-1vcpu-1gb up to mars-16vcpu-32gb.

The agent inside it was driven by DigitalOcean's own inference endpoint rather than a third-party key. A single prompt through DeepSeek v4 Pro came back in 7.3 seconds, 15,793 tokens in and 91 out, having written the file I asked for. The inference and the sandbox billing arrive on one account, which is a smaller convenience than the pause thing but not nothing.

Forking, which is where it got strange

There is a fork command, and I assumed it did what checkpoint-and-restore products usually do, which is snapshot the disk and give you a second sandbox with the same files.

It turns out not to. I forked the running parent twice, took 30.98 seconds for both, and then checked the ticker process in all three sandboxes.

ticker PID counter shortly after counter a minute later
parent 590 177 244
child 1 590 165 231
child 2 590 162 228

Every one of them had the same process at the same PID, counting, from the value it held at the moment the fork was taken. Three copies of one running program, diverging from a common ancestor. If you have ever wanted to run an agent up to a decision point and then explore four different choices from exactly that state, without replaying the work that got you there, this is the primitive that does it.

Checkpointing separately took 25.25 seconds and reported a size of roughly 111 GB, which is the sparse allocation rather than anything you are storing.

Collecting the wall-clock cost of every operation in one place, since the spread between them is the thing that would shape how you use it:

Measured duration of each session operation

Creating a session is the expensive one at 15.97 seconds. Pausing and resuming are close enough to instant that you would not build around them, which is what makes the pause worth reaching for in the first place.

Two things I would want stated plainly

Egress from a fresh sandbox is open, and I confirmed that by reaching both example.com and api.github.com from one without configuring anything at all. Adding a single host to the manifest flips the behaviour entirely:

egress:
  - api.github.com
Enter fullscreen mode Exit fullscreen mode

After that, api.github.com returned 200 and everything else returned nothing at all, which is the right design, since naming one host is an unambiguous statement that you want a deny-by-default posture. But an unconfigured session has a general-purpose language model with a shell and the open internet, and the default is the permissive one.

The other thing is that HARNESS_INFERENCE_API_KEY is readable as an ordinary environment variable from inside the sandbox, so anything running in there can print it. That is unavoidable if the process is going to call the inference endpoint, and the same is true of every runtime I know of, but it is the reason the egress allowlist matters more than it looks.

And this is a public preview in a single region without an SLA. Everything above is true of what shipped and none of it is a commitment about what will ship.

What I got wrong

Twice, and both times the same shape of mistake.

The first attempt at starting the ticker used bash -s < piped through the exec channel. The command reported success, the log file appeared with a few lines in it, and then it stopped, because the process died with the channel. I spent a while reading pause documentation for an answer to a problem that was not about pausing at all.

The second one was worse, because it produced a plausible number. I checked whether the checkpoint flag existed by running checkpoint create --name, got an error, and nearly wrote that checkpoints could not be labelled. The flag is --label. Three commands in this session refused an argument I had assumed from the shape of other tools, and in each case the refusal text was the thing that told me, which is an argument for quoting error output rather than paraphrasing it.

What it cost

Four sessions, two of them forks, one checkpoint of a 111 GB sparse image, a few dozen exec calls and one inference request. The account balance went from $4.40 to $4.33, so the whole morning came to seven cents.

I mention the figure because the thing that usually stops people testing a runtime properly is the fear of leaving something running, and at these prices the honest advice is to go and try it yourself rather than trust my numbers.

The manifests, the raw tick log and both charts are in a small repo if you want to reproduce it. Three commands is the whole of it:

doctl harness-runtime create -f agent.yaml
doctl harness-runtime pause <session-id>
doctl harness-runtime resume <session-id>
Enter fullscreen mode Exit fullscreen mode

What to take from it

If you are evaluating any agent runtime, the pause claim is the one to test first and the one nobody tests, because the naive version of the test passes trivially. Put something in memory that is never written down, and see whether it is still counting on the other side.

And if it is, the operational consequences are larger than the feature description suggests. An agent that costs nothing while it is paused can wait for a human review instead of being torn down and rebuilt. An agent you can fork from a running state can be tried three ways from one expensive setup. Both of those change how you would structure a long-running job, and neither of them is the sort of thing you find out from a pricing page.

Top comments (0)