Building payment workflows that remain correct when events are duplicated, delayed, missing, or delivered out of order.
Payment integrations often look simple in a sequence diagram:
- Create a transaction.
- Send it to a payment provider.
- Receive a webhook.
- Mark the transaction successful.
The difficult part begins when production stops following the diagram.
A webhook may arrive twice. A “processing” event may arrive after a “successful” event. The provider may accept an API request with HTTP 200 while the underlying payment remains pending. A status check may return a transport-level success code that has nothing to do with the financial result. In some cases, the final webhook may never arrive.
When money is involved, a webhook cannot be treated as a command to change a balance. It is evidence about an external event. The local transaction record must decide whether that evidence permits a state transition and whether the related financial effect has already happened.
Store the transaction before trusting asynchronous events
Every external payment should map to an internal transaction created before the asynchronous lifecycle begins. That record should contain enough stable identifiers to match later provider events without relying on mutable fields such as a description or customer-entered name.
A simplified record might contain:
class Payment:
reference: str
provider_reference: str | None
status: str
amount: Decimal
currency: str
credited_at: datetime | None
The local reference identifies the business operation. The provider reference identifies the external attempt. They serve different purposes and should not be silently substituted for each other.
Separate transport success from financial success
An HTTP success response can mean:
- the provider received the request;
- the request passed initial validation;
- the operation was queued;
- the provider returned a status payload successfully; or
- the financial transaction completed.
Only the last meaning should trigger a final credit or completion.
I prefer translating provider responses into a small internal vocabulary:
PENDING = {"accepted", "queued", "processing"}
SUCCESS = {"settled", "completed"}
FAILURE = {"failed", "rejected", "cancelled"}
The real mapping varies by provider, but the rule remains constant: unknown and intermediate states should not be promoted to success.
This is especially important when a provider exposes both a general response code and a transaction status. A successful status-check request does not prove that the payment succeeded.
Make every financial effect idempotent
Webhook delivery is normally at least once. Therefore, the same success event must be safe to process repeatedly.
Checking only the transaction status is helpful but incomplete. Two workers can read the same pending status before either one commits. Both may then attempt the credit.
A stronger design protects the financial effect with its own deterministic identity:
with database_transaction():
payment = lock_payment(reference)
if ledger_entry_exists(key=f"deposit:{payment.reference}"):
return "already processed"
if payment.status in TERMINAL_STATES:
return "ignored"
create_ledger_credit(
key=f"deposit:{payment.reference}",
amount=payment.amount,
currency=payment.currency,
)
payment.status = "completed"
payment.save()
The row lock prevents concurrent transitions. The unique ledger key prevents duplicate financial effects even if application logic is retried.
Protect terminal states from older events
External systems do not always deliver events in chronological order. Once a transaction reaches a trusted terminal state, older events must not move it backward.
For example:
ALLOWED_TRANSITIONS = {
"pending": {"processing", "completed", "failed"},
"processing": {"completed", "failed"},
"completed": set(),
"failed": set(),
}
This prevents a delayed “processing” notification from reopening a completed transaction. It also stops a late failure from automatically reversing a payment that has already been independently confirmed, unless the product explicitly supports a separate reversal workflow.
Preserve raw evidence without exposing it to customers
Webhook payloads are useful for investigation, reconciliation, and provider support. Retain them with appropriate access controls, but don't make them the customer-facing transaction model.
Store both:
- a normalized internal status and safe customer message;
- the original provider event for staff diagnostics.
This separation prevents technical provider errors, internal account balances, or implementation details from leaking into the customer experience.
Add an active recovery path
Even a perfect webhook handler cannot process an event that never arrives. Transactions that remain in a non-terminal state need background reconciliation.
A recovery worker can:
- Find processing transactions older than a safe threshold.
- Query the provider using the recorded provider reference.
- Normalize the returned state.
- Apply the same guarded transition used by webhook processing.
- Retry with limits when the provider remains inconclusive.
The webhook and reconciliation paths should converge on the same domain function. Otherwise, each path develops different rules for crediting, failure, and refunds.
The design principle
The transaction record is authoritative for what the platform has applied. The provider remains authoritative for what happened on its rail. Webhooks and status checks are two ways of collecting provider evidence; neither should bypass local invariants.
That distinction turns webhook handling from a collection of callbacks into a reliable payment lifecycle.
The practical outcome of this architecture is predictable behavior under uncertainty. Duplicate delivery becomes harmless, intermediate states remain intermediate, and a missing callback becomes a reconciliation problem instead of an indefinitely stuck payment. Those are properties that can be verified through invariants and tests without relying on an unsupported performance claim.
Top comments (0)