Designing retry-safe financial operations across clients, workers, databases, and external providers.
Payment clients retry requests. Load balancers retry requests. Background workers retry tasks. Users press a button again when the screen appears stuck. Providers may time out after accepting an operation, leaving the caller unsure whether it succeeded.
Retries are normal. Duplicate financial effects are not.
An idempotency header is useful, but it is only the beginning of an idempotency design.
Define the operation being protected
An idempotency key needs a scope. The same string may be valid for one operation and invalid for another.
A practical internal record can include:
class IdempotencyRecord:
operation: str
owner_key: str
idempotency_key: UUID
request_hash: str
completed: bool
resource_id: str | None
The pair (operation, owner_key) identifies the business attempt being protected. Examples might include creating one external account for a user and currency, submitting one payout request, or provisioning one wallet on a network.
The external idempotency key is then stable for retries of that same attempt.
Bind the key to the request body
Reusing an idempotency key with a different payload is not a harmless retry.
Imagine the first request attempts to send 100 units and the second attempts to send 1,000 using the same key. Returning the result of the first request would surprise the caller. Sending the second request with the same external key may be rejected by the provider. Executing both defeats idempotency entirely.
Canonicalize and hash the request:
def request_hash(payload):
canonical = json.dumps(
payload,
sort_keys=True,
separators=(",", ":"),
default=str,
)
return sha256(canonical.encode()).hexdigest()
When a retry arrives:
- same operation, owner, and hash: reuse the existing attempt;
- same operation and owner but different hash while pending: reject the conflict;
- completed attempt: return the stored result or begin a clearly defined new operation.
Lock creation of the idempotency record
Two workers can attempt to create the same record simultaneously. The application therefore needs both a database uniqueness constraint and transactional locking.
with database_transaction():
record, created = (
IdempotencyRecord.objects
.select_for_update()
.get_or_create(
operation=operation,
owner_key=owner_key,
defaults={"request_hash": request_hash(payload)},
)
)
The database constraint is the final arbiter. Code-level checks improve behavior and error messages, but they do not replace uniqueness.
Persist the resulting external resource
An idempotency record is more useful when it stores the resource created by the operation.
Suppose the external request succeeds but the API response is lost. On retry, the application should be able to determine whether an external resource was already linked locally rather than blindly creating another one.
The completed record can retain:
- the local resource identifier;
- the external resource identifier;
- a sanitized response summary; and
- completion time.
This allows safe replay and operational investigation.
Protect the ledger independently
Provider request idempotency does not automatically make the local financial effect idempotent.
The external transfer may be created once while its success webhook is delivered multiple times. Therefore, use separate idempotency boundaries:
- Command idempotency prevents duplicate provider operations.
- Event idempotency prevents processing the same provider event repeatedly.
- Ledger idempotency prevents applying the same credit or debit twice.
- Notification idempotency prevents duplicate customer messages.
These keys may derive from the same business reference, but they protect different effects.
Decide what completion means
An operation is not always complete when the provider accepts it.
For asynchronous payments, useful stages include:
- request safely submitted;
- external resource created;
- financial operation processing;
- terminal result received; and
- local ledger effect applied.
The idempotency record should reflect the boundary it protects. Marking everything completed immediately after an HTTP response can make later recovery difficult.
Do not use random keys without durable storage
Generating a new UUID for every retry provides uniqueness, not idempotency. The caller must reuse the same key for the same logical operation, or the server must derive and persist a stable association.
The key’s lifecycle matters more than its randomness.
The design principle
Idempotency is a business guarantee that repeated execution produces one intended effect. The HTTP header is simply one mechanism for carrying that intent across a network boundary.
A reliable implementation defines the operation’s scope, binds it to a canonical request, serializes competing attempts, stores the result, and independently protects downstream financial effects.
The resulting behavior is deterministic even when the network is not. A timeout can be retried without creating a new business operation, a conflicting payload is rejected, and downstream effects remain independently protected. This is a stronger guarantee than assuming retries will be rare.
Top comments (2)
The (operation, owner_key) scoping plus hashing the canonical payload is the part most teams skip — we learned the hard way that reusing a key with a different amount has to be a hard conflict, not a silent retry. Your split between command, event, and ledger idempotency also maps exactly to why a provider ack can't be treated as 'money moved' until the local ledger effect is recorded once.
Exactly. The key identifies the business operation, while the payload hash confirms it is still the same request. If the amount or destination changes, treating it as a retry creates ambiguity.
And yes, provider acceptance and local settlement are separate guarantees. Command, event, and ledger idempotency each protect a different boundary; none can safely replace the others.