Temporal's promise is durable execution. You write a workflow as ordinary code, and if the process running it crashes halfway through, a new process picks it up and carries on from where it stopped, as though nothing had happened. I wanted to see that promise under pressure, and in particular I wanted to know what "as though nothing had happened" means for the things a workflow does to the outside world.
So I ran a ten-step workflow and killed the worker running it eight times, with kill -9, at random moments. The workflow finished exactly once, as promised, and its first step ran four times, which is also as documented if you read closely enough.
The setup
Everything here runs locally on the Temporal CLI's development server, version 1.9.1 bundling Server 1.32.0, with the Python SDK 1.34.0.
The workflow is deliberately plain: ten activities in sequence, each called step. Every activity writes a line to a file before it does anything else, standing in for a real side effect such as charging a card or sending an email, and then spends three seconds on slow work:
@activity.defn
async def step(order: str, n: int) -> str:
with open(SIDEFX, "a") as f:
f.write(f"... step={n} attempt={activity.info().attempt} pid={os.getpid()}\n")
await asyncio.sleep(STEP_SECONDS)
return f"step {n} done"
Writing first and working second is the honest order for this test, because it is the order real code tends to use: you call the payment provider, then you update your own records. A crash between the two is exactly the case worth knowing about.
A small driver starts the workflow, then repeatedly sleeps between two and six seconds, sends SIGKILL to the worker, and immediately starts a fresh one. Run without interruptions, the same workflow takes thirty seconds from start to finish.
One workflow, four charges
After eight kills, it reported completed, after 69.2 seconds rather than thirty:
The result is exactly what the workflow should have returned, each step's output once and in order. The workflow itself was not confused at all. Then I counted the lines in the side-effect file:
step 1 ran 4 times
step 2 ran 1 time
...
step 10 ran 1 time
Temporal promises that the workflow completes exactly once and that each activity's result is recorded once. It does not promise that each activity runs once, and it says so in the documentation, which describes activities as at-least-once. Knowing that in principle is different from watching one step happen four times while the dashboard reports a single clean success.
Why it took eight seconds to notice
Lining the kills up against the side effects explains the whole run:
The first kill landed 2.9 seconds in, while step 1 was doing its slow work. A new worker was running a fraction of a second later, and then nothing happened for eight seconds.
The server has no way to see a process die. As far as it knows, step 1 is still running on the old worker, and the only thing that will change its mind is the activity's start_to_close_timeout expiring, which I had set to ten seconds. The retries landed at precisely that timeout plus the retry policy's backoff:
| started | gap since the previous attempt | |
|---|---|---|
| attempt 1 | 0.3 s | |
| attempt 2 | 11.3 s | 10 s timeout + 1 s |
| attempt 3 | 23.3 s | 10 s timeout + 2 s |
| attempt 4 | 37.3 s | 10 s timeout + 4 s |
The other five kills changed nothing at all. They hit a worker that was idle, sitting there waiting for the server's timeout to expire on a task that a dead process still nominally held.
The history doesn't show it either
If you went looking for those duplicates in Temporal's own record, you would struggle. The event history holds one ActivityTaskStarted event per step, ten for ten steps, and the one for step 1 reads:
ACTIVITY_TASK_STARTED attempt=4 lastFailure=activity StartToClose timeout
The three failed attempts are not stored as events of their own. They are folded into a counter on the attempt that eventually succeeded. Read the history and step 1 looks like a single execution that happened to need four tries; read the side-effect log and it ran four times, and both are telling the truth about different things.
Heartbeating made the duplicates worse
The standard advice for a slow-to-notice crash is heartbeating. The activity reports in regularly, and the server declares it dead after a short silence rather than waiting out the whole timeout. I added a heartbeat every half second with a two-second heartbeat timeout, and reran with the identical kill schedule.
Detection got much faster, a kill at 2.9 seconds now retried at 4.9. The duplicates got worse. To make sure that was not luck, I ran four different kill schedules in both modes:
| duplicate side effects | kills that hit an idle worker | time to finish | |
|---|---|---|---|
| timeout only, 10 s | 3, 3, 2, 2 | 5, 5, 6, 6 | median 66.5 s |
| heartbeat, 2 s | 6, 4, 5, 5 | 2, 4, 4, 3 | median 70.7 s |
Heartbeating doubled the duplicates and finished no faster. The reason is in the middle column. With a slow timeout, most kills land while the replacement worker is idle, waiting for the server to give up on the old one. With fast detection, the retry is already running when the next kill arrives, so far more of the kills interrupt real work. And because every failure doubles the backoff, the time saved by detecting faster gets spent waiting for the next retry.
I want to be fair to heartbeating here, because this is a pathological test. A crash every two to six seconds is not a production workload. When crashes are rare, cutting detection from ten seconds to two is a real improvement and you should take it. What the experiment shows is narrower and more important: anything that makes retries faster also makes them more frequent, and every retry runs your side effect again.
The fix is a key, not a setting
No retry configuration solves this, because the problem is in the activity, not the policy. The activity has to be safe to run more than once.
Temporal gives you what you need for that. activity.info() includes the workflow ID and the activity's ID, and the combination is identical on every retry of the same step. I used it as an idempotency key against a table with a unique constraint:
key = f"{info.workflow_id}:{info.activity_id}"
applied = con.execute("insert or ignore into charges values (?, ?)",
(key, time.time())).rowcount == 1
Then I reran the worst case, heartbeating on the kill schedule that had run one step six times:
step 1: executed 6x, charged 1x
...
TOTAL : 15 executions, 10 charges
The retries still happened, but now they did no harm. With a real payment provider you would pass the same key to their idempotency header rather than keeping your own table, but the principle is identical.
What the determinism check actually checks
The other half of Temporal's contract is that workflow code must be deterministic. When a worker picks up a workflow, it re-runs the code from the beginning against the recorded history, and the code has to make the same decisions it made the first time. I wanted to know how strict that check is, so I took the real 65-event history from the eight-kill run and replayed it against deliberately altered versions of the workflow:
| change to the workflow code | replay |
|---|---|
| unchanged | accepted |
| a new activity inserted before step 1 | rejected |
| the loop runs nine steps instead of ten | rejected |
| the same ten steps in reverse order | accepted |
| completely different arguments to every step | accepted |
| the timeout changed from 10 s to 999 s | accepted |
The rejections come with clear messages:
[TMPRL1100] Nondeterminism error: Activity type of scheduled event 'step'
does not match activity type of activity command 'validate'
The accepted cases are the interesting ones. Running the steps backwards replayed cleanly. So did passing every step a different order ID and a different step number. The check compares the shape of what the workflow does, which kinds of activity it schedules and in what order, not the arguments it passes them.
That also explains why the idempotency key worked. The activity IDs Temporal assigned were simply 1 to 10, each step's position in the workflow, so the reversed workflow's first activity was matched to the original's first activity regardless of what it actually did. One consequence follows that I did not test directly: if you reorder steps in a new version of a workflow while using position-based keys, a key can end up attached to a different step than the one that originally used it.
This was the Python SDK. Other SDKs share much of the same core but I have not checked whether they behave identically, so treat the table as describing 1.34.0 specifically.
What I got wrong
The first version of the workflow never ran. I read the timeout from an environment variable inside workflow code, and the Python SDK's sandbox refused at runtime:
RestrictedWorkflowAccessError: Cannot access os.environ.get from inside a workflow.
That is the sandbox doing its job, since an environment variable can differ between the original run and a replay, and the timeout now arrives as a workflow argument instead, where it is recorded in the history like everything else.
My first replay test reported every case as rejected, including the unchanged code, which was there as a control. The sandbox re-imports the module that defines the workflow, and I had defined the workflow variants in the same file that called asyncio.run(main()), so the re-import tried to run the whole test again. Moving the variants into their own module fixed it. Without the control I would have published a table saying Temporal rejects everything.
One background run looked stuck for long enough that I went looking for a bug. It was simply slow: each kill during a step costs the full timeout, and the backoff doubles each time.
And twice in this experiment I killed my own shell. To stop the dev server I ran pkill -f "temporal server start-dev", which matched the command line of the shell running it. Writing the pattern as [t]emporal matches the server's command line but not its own, and I now use it everywhere.
What to take from it
Temporal did exactly what it says. Eight hard kills, no lost progress, no corrupted state, and the workflow's result was correct. The distinction worth carrying away is between the workflow, which really does complete once, and the activities inside it, which run at least once and in this test ran up to six times.
So write every activity as though it will run twice, because under failure it will. Give side effects an idempotency key built from the workflow ID and activity ID, and pass it on to anything external that accepts one. Set heartbeat timeouts for faster recovery, but understand that faster retries mean more retries, which makes the key more important, not less.
And treat a passing replay as weaker evidence than it sounds. It proves the workflow still takes the same shape, not that it still does the same thing.
The workflow, the kill driver, the raw side-effect log and the exported history are in a small repository. The replay test reads the recorded history from disk, so it runs without a Temporal server.



Top comments (2)
tr.ee/dev-to
Really thoughtful post! Documenting real-world engineering hurdles and actionable solutions like this brings immense value to the community.