Introduction
Durable execution guarantees that your workflow finishes. It does not guarantee that each step runs only once. Picture a workflow that charges a credit card for 49 EUR. The payment provider accepts the charge, and then the pod is evicted before the workflow records the result. On restart, the workflow sees an unfinished step and charges the card again. Your customer now pays 98 EUR.
If you are new to durable execution, this can feel like a bug in the engine. It is not. Durable execution engines, including Dapr Workflow, run activities at least once, and making a side effect happen exactly once is the job of your activity code. In this post, you'll see why the engine behaves this way, how to spot the steps that need protection, and how to use the task execution ID in a C# activity as an idempotency key.
Durable Execution Guarantees Completion, Not Single Execution
A Dapr Workflow app has two kinds of code. The workflow code decides what happens next: call this activity, wait for that event, branch on a result. The activity code does the actual work that touches the outside world, such as calling a payment provider, writing to a database, or sending an email. The workflow engine persists every step, so when your app crashes it can pick up where it left off instead of starting over. If you want a primer on this model, watch this session on durable execution with Dapr.
There is a catch in "pick up where it left off". The engine can only store that a step happened after the step has happened. Between the moment your activity charges the card and the moment the engine writes down the result, there is always a window. If the process dies inside that window, the engine has no record of the charge, so the only safe choice is to run the activity again.
That is why durable execution gives you completion, not exactly-once execution. Exactly-once side effects come from making each activity idempotent: running it twice with the same input has the same effect as running it once. This is not specific to Dapr. Temporal and Azure Durable Functions make the same trade-off and give the same advice.
How Dapr Workflow Records Progress
Here is the order workflow from the intro, written with the Dapr .NET SDK:
public class OrderWorkflow : Workflow<Order, OrderResult>
{
public override async Task<OrderResult> RunAsync(WorkflowContext context, Order order)
{
var payment = await context.CallActivityAsync<PaymentResult>(
nameof(ChargeCardActivity),
new ChargeRequest(order.CustomerId, order.Amount));
await context.CallActivityAsync(nameof(ShipOrderActivity), order);
return new OrderResult(payment.PaymentId);
}
}
The workflow code does not call ChargeCardActivity directly. The orchestration lifecycle describes what happens instead. The Dapr sidecar sends your app an OrchestratorWorkItem over a long-lived GetWorkItems gRPC stream. That work item contains the full history of the workflow instance. The SDK replays your RunAsync method from the top against that history. When it reaches CallActivityAsync, it looks for a recorded result. If there is one, it returns the recorded result without running the activity. If there isn't, it emits a ScheduleTask action and suspends the workflow.
The engine then appends a TaskScheduled event to the history and dispatches the activity. When the activity reports back, the engine records a TaskCompleted (or TaskFailed) event, and the next replay picks up the result.
According to the state and history docs, this history is the source of truth. It is stored in your configured state store as a list of serialized events, with keys like wf-history-<instance_id>-<index>. For the order workflow, the history at the moment of the pod eviction looks like this:
ExecutionStarted OrderWorkflow
TaskScheduled ChargeCardActivity
<- pod evicted here, card already charged
There is no TaskCompleted for ChargeCardActivity. As far as the engine knows, the charge never finished.
Why Activities Run at Least Once
The activity lifecycle shows where the window sits. When the engine dispatches an activity, the sidecar sends an ActivityWorkItem over the same GetWorkItems stream. The work item carries the activity name, its input, the workflow instance_id, a task_id, a task_execution_id, and a completion_token. The SDK runs your activity code and then calls CompleteActivityTask with the result (or failure_details) and the completion_token, as defined in the execution API.
Your side effect happens while the activity runs. The engine learns about it only when CompleteActivityTask arrives. Anything that breaks the chain between those two points, such as a pod eviction, an out-of-memory kill, a node failure, or a dropped network connection, leaves a TaskScheduled event without a matching TaskCompleted. The activity lifecycle doc names this case directly: activities "might be executed more than once (e.g., if the worker crashes after execution but before reporting completion)".
Retry policies are a second source of repeats. If an activity fails and the workflow has a retry policy, the engine schedules the activity again with the same name and input. A failure from your activity's point of view is not always a failure on the other side. If the HTTP call to the payment provider times out, your activity throws, yet the provider may have processed the charge anyway.
The third case is a slow worker. If a worker takes too long and the task is handed to another worker, the original one can still finish. The sidecar uses the completion_token to ignore the late response, but by then the side effect has already happened.
In all three cases, the engine protects its own bookkeeping, not your side effects. That is why at-least-once is the norm for durable execution engines.
How to Spot the Steps That Need Idempotency
Not every activity needs extra work. For each activity in your workflow, ask one question: if this runs twice with the same input, is the outcome different from running it once?
Activities that usually need protection:
- Charging, refunding, or transferring money
- Sending emails, text messages, or push notifications
- Inserting records with generated IDs, where a second insert creates a duplicate row
- Publishing messages to a broker or event stream
- Incrementing counters or adjusting stock levels
- Calling an LLM tool or external API that acts on another system, such as creating a ticket
Activities that are usually safe already:
- Pure computation on the input
- Reads, such as GET requests and database queries
- Upserts keyed on a natural ID, such as the order ID
- Setting a value to a fixed state, such as "set order status to Shipped"
Don't forget compensation activities. If your workflow uses the saga pattern, the refund that undoes a charge can run twice as well, and refunding a customer twice is just as wrong as charging them twice.
Keep this separate from workflow determinism. Determinism is about your workflow code producing the same decisions on every replay, so no DateTime.UtcNow, random numbers, or direct I/O in RunAsync. Idempotency is about your activity code having the same effect when it runs more than once. You need both, but they solve different problems.
Using the Task Execution ID as an Idempotency Key
The standard fix is an idempotency key. You send a key with the request, and the receiving service remembers which keys it has already processed. When the same key arrives a second time, the service returns the original result instead of doing the work again. Most payment providers support this through an Idempotency-Key HTTP header.
The key has to be the same every time the activity runs for the same step, and different for every other step. Since Dapr .NET SDK v1.17, WorkflowActivityContext exposes a TaskExecutionId property that fits this. The ID is created when the workflow schedules the activity and is stored with the TaskScheduled event in the history. When the engine dispatches that scheduled task again after a crash, a pod eviction, or a worker timeout, your activity receives the same ID. Treat it as an opaque string and don't parse it.
Here is the payment activity with the task execution ID passed as the idempotency key:
public class ChargeCardActivity(
IHttpClientFactory httpClientFactory,
ILogger<ChargeCardActivity> logger)
: WorkflowActivity<ChargeRequest, PaymentResult>
{
public override async Task<PaymentResult> RunAsync(
WorkflowActivityContext context, ChargeRequest request)
{
var idempotencyKey = context.TaskExecutionId;
logger.LogInformation(
"Charging {Amount} EUR for {CustomerId} with idempotency key {Key}",
request.Amount, request.CustomerId, idempotencyKey);
var httpClient = httpClientFactory.CreateClient("payments");
using var message = new HttpRequestMessage(HttpMethod.Post, "/v1/payments")
{
Content = JsonContent.Create(request)
};
message.Headers.Add("Idempotency-Key", idempotencyKey);
using var response = await httpClient.SendAsync(message);
response.EnsureSuccessStatusCode();
return (await response.Content.ReadFromJsonAsync<PaymentResult>())!;
}
}
Replay the eviction scenario with this version:
- The engine dispatches
ChargeCardActivitywith task execution IDa1b2.... - The provider charges 49 EUR and stores the payment under key
a1b2.... - The pod is evicted before
CompleteActivityTaskreaches the sidecar. - The workflow restarts, and the engine dispatches the same scheduled task again with task execution ID
a1b2.... - The provider recognizes the key and returns the original payment without charging the card.
- The engine records
TaskCompleted, and the workflow moves on to shipping.
Your customer pays 49 EUR. The same approach works for any service that accepts a deduplication key, such as email providers, message brokers that support message deduplication IDs, and your own APIs.
When the Idempotency Key Expires
An idempotency key only helps while the service still remembers it. Many providers clear keys after a fixed period. Stripe, for example, can remove keys once they are at least 24 hours old. Check the documentation of every service you call for its retention period.
Durable workflows can run for days or weeks, and a second dispatch doesn't always follow the first one quickly. An outage, a long retry backoff, or a suspended workflow that resumes later can all push the second attempt past the retention period. At that point the provider treats the request as new. Some services don't support idempotency keys at all.
In those cases you need a record you own: a dedupe table in your own database, keyed on the task execution ID. Before doing the work, the activity checks the table. After doing the work, it stores the result.
public class SendInvoiceActivity(InvoiceDbContext db, IEmailSender emailSender)
: WorkflowActivity<InvoiceRequest, InvoiceResult>
{
public override async Task<InvoiceResult> RunAsync(
WorkflowActivityContext context, InvoiceRequest request)
{
var key = context.TaskExecutionId;
var processed = await db.ProcessedTasks.FindAsync(key);
if (processed is not null)
{
return new InvoiceResult(processed.MessageId);
}
var messageId = await emailSender.SendAsync(
request.To, request.Subject, request.Body);
db.ProcessedTasks.Add(new ProcessedTask(key, messageId, DateTimeOffset.UtcNow));
await db.SaveChangesAsync();
return new InvoiceResult(messageId);
}
}
Make the task execution ID the primary key of ProcessedTasks. A unique constraint stops two concurrent attempts, such as a slow worker and its replacement, from both recording a result. Keep rows for longer than your longest-running workflow, and clean up older rows with a scheduled job.
This version still has a gap. If the process dies between SendAsync and SaveChangesAsync, the email goes out twice. To close it, insert a Pending row before calling the external service and mark it Completed afterwards. When an attempt finds a Pending row, it knows a previous attempt may have done the work, and it can check with the provider before trying again. Where the provider supports keys, send the task execution ID to it as well: the provider's key covers the short window, and your table covers everything after the key expires. For more on choosing between keys and this kind of reconciliation, see the Diagrid FAQ on idempotency keys and reconciliation logic.
Trade-offs and Considerations
Check how the task execution ID behaves with retry policies before you rely on it. The .NET SDK documentation describes it as stable across retries, while the activity lifecycle protocol doc says it changes every time the engine retries the activity. For the payment timeout case, that difference decides whether the provider sees one charge or two. Test it on your Dapr and SDK versions. If you need one key across all retry attempts, build it in the workflow from the instance ID and a step name, for example $"{context.InstanceId}:charge-card", and pass it as activity input. Use a different step name for every call, including calls inside a loop.
Scope each key to one side effect. If an activity both charges a card and sends a receipt, it needs two keys, such as {TaskExecutionId}:charge and {TaskExecutionId}:receipt. Splitting the work into two activities is usually cleaner.
Idempotency has a cost. A dedupe table adds a database round trip to every call and a cleanup job to your operations work. Spend that effort on the activities from the checklist above, not on every activity.
Don't move side effects into the workflow code to avoid retries. Workflow code replays on every step, so a side effect there runs far more often, and it breaks determinism.
Finally, log the key in every activity. Idempotency stops the duplicate from reaching your customer, and the logs tell you how often it happened.
Summary
Durable execution keeps your workflow moving through crashes and pod evictions, but it does so by running activities at least once. The Dapr workflow protocol makes the reason visible: the engine only knows an activity finished once a TaskCompleted event lands in the history, so any failure in between leads to a second dispatch. The fix belongs in your activities. Find the steps with side effects, pass the task execution ID as an idempotency key to the services that support it, and back it with a dedupe table when the downstream service forgets keys too soon or doesn't support them at all. Your 49 EUR charge then stays 49 EUR.
To build a solid foundation in the Dapr Workflow programming model, including activity retries and error handling, take the free Dapr Workflow track on Dapr University. If you prefer video, this session on durable execution with Dapr covers how replay and history work.
Top comments (0)