Distributed transactions become interesting exactly when the happy path stops being reliable.
A basic recharge implementation might look like this:
charge card
↓
call recharge provider
↓
return success
In a local test environment, that can appear perfectly adequate.
Production introduces:
- duplicate requests;
- timeouts;
- delayed provider responses;
- asynchronous status updates;
- partial failures;
- process crashes;
- payment success followed by fulfilment uncertainty;
- users refreshing the page;
- workers retrying jobs.
The underlying prepaid mobile top-up use case is a useful example because a single user action can involve payment, external fulfilment, asynchronous confirmation and an operation that should never be executed twice accidentally.
The solution is not to add more try/catch blocks.
The transaction needs an explicit lifecycle.
Start with a transaction record before external execution
Do not wait until the provider responds successfully before creating a database record.
Create the transaction first.
For example:
{
"id": "txn_123",
"status": "created",
"recipient": "+447700900123",
"quote_id": "q_456",
"idempotency_key": "abc-123",
"created_at": "..."
}
Now every later operation has an internal identity.
That identity can be used in:
- logs;
- payment metadata;
- provider metadata;
- support tools;
- reconciliation jobs.
A transaction that fails midway still exists.
That is valuable.
Use a state machine instead of a success flag
A boolean cannot represent asynchronous processing well.
This:
success = null
quickly becomes ambiguous.
A state model is clearer.
For example:
created
payment_pending
payment_authorized
submitted
processing
succeeded
failed
Depending on the architecture, you might also need:
payment_failed
cancelled
refunded
manual_review
The exact state names matter less than defining what each one means.
For every state, ask:
- What must already have happened?
- Which transitions are allowed?
- Is the state terminal?
- Can a worker safely retry from here?
Idempotency belongs at the transaction boundary
Imagine this request:
POST /api/recharges
The user double-clicks.
Or the frontend retries because the first response takes too long.
Without protection:
Request 1 → recharge
Request 2 → recharge
The customer may pay twice or the recipient may receive duplicate value.
Instead, require an idempotency key:
Idempotency-Key: 70356c3a-...
Store it with a unique database constraint.
Pseudo-code:
existing = find_transaction(idempotency_key)
if existing:
return existing
transaction = create_transaction(idempotency_key)
process(transaction)
The unique constraint matters because two concurrent application workers can otherwise both pass an application-level “does this exist?” check.
Correctness should survive concurrency.
Do not blindly retry an ambiguous provider request
Suppose you send:
POST /provider/recharge
The provider completes the recharge.
Your connection drops before the HTTP response arrives.
From your application's perspective:
request timeout
From the provider's perspective:
recharge succeeded
If your retry policy says:
timeout → retry immediately
you may create a duplicate.
A timeout means:
We do not know the outcome.
That is different from:
The operation failed.
If the provider supports idempotency, use it.
If it provides transaction lookup, query the original request.
If it sends asynchronous callbacks, wait for the callback within a reasonable reconciliation window.
Retries need business context.
Give every provider operation a correlation identifier
A useful integration keeps both sides traceable.
For example:
internal transaction:
txn_123
provider request reference:
mobiletopup-txn_123
provider transaction:
ext_987
Store the provider transaction ID as soon as you receive it.
Then support and reconciliation processes can move between:
your system ↔ provider system
without searching by amount and timestamp.
Correlation IDs are cheap.
Debugging without them is not.
Separate payment state from recharge state
One of the most important modeling decisions is not to treat payment success as transaction success.
Consider:
Payment: authorized
Recharge: processing
That is a perfectly valid intermediate state.
Likewise:
Payment: failed
Recharge: not submitted
is very different from:
Payment: captured
Recharge: provider rejected
A transaction may therefore contain separate dimensions:
{
"payment_status": "captured",
"fulfilment_status": "processing"
}
The overall user-facing status can be derived from those states.
Do not destroy useful information by compressing both into one column too early.
Make asynchronous completion normal
External fulfilment often works better as an asynchronous workflow.
A simplified architecture:
API request
↓
create transaction
↓
authorize/capture payment
↓
enqueue recharge job
↓
provider submission
↓
processing
↓
webhook or polling
↓
final state
The API does not need to keep an HTTP connection open for the entire provider lifecycle.
The frontend can poll:
GET /api/recharges/txn_123
or receive updates through another mechanism.
This also makes temporary provider slowness easier to absorb.
Webhooks need the same defensive engineering
A webhook can arrive:
- once;
- twice;
- out of order;
- much later than expected.
Treat webhook processing as idempotent.
For example:
provider event ID
+
unique constraint
If the same event arrives twice, processing it twice should not produce duplicate side effects.
Also validate the webhook's authenticity using the provider's supported verification mechanism.
Do not trust an arbitrary public POST request simply because it contains a transaction ID.
Add reconciliation even if webhooks exist
Webhooks are useful.
They are not magic.
Events can be delayed or lost because of:
- endpoint downtime;
- configuration mistakes;
- provider incidents;
- networking problems;
- internal processing bugs.
A reconciliation worker can periodically find transactions such as:
status = processing
AND updated_at < now - threshold
and query the provider.
This gives the system a recovery path that does not depend on every asynchronous event arriving perfectly.
Store events, not only the latest state
A row that currently says:
status = succeeded
does not explain how it got there.
An event history might say:
12:00:01 transaction_created
12:00:05 payment_authorized
12:00:06 provider_submitted
12:00:07 provider_acknowledged
12:00:24 provider_processing
12:00:32 recharge_succeeded
That timeline is enormously useful.
You can implement it with:
- a dedicated transaction-events table;
- structured application events;
- or an appropriate event-sourcing approach where warranted.
You do not need full event sourcing merely to maintain an audit trail.
Retries should be state-aware
A retry policy such as:
retry every exception three times
is dangerous for financial or fulfilment operations.
Instead, decide by operation type.
A read-only provider lookup?
Usually safe to retry.
Submitting a recharge with provider idempotency?
Potentially safe to retry using the same key.
Submitting without provider idempotency after an ambiguous timeout?
Reconcile first.
Retry logic belongs to the business workflow, not just the HTTP client.
Design for support from day one
Operational support will eventually need answers to questions such as:
- Did the payment complete?
- Was the recharge submitted?
- What external ID did the provider return?
- When was the last provider status check?
- Did a webhook arrive?
- Was the transaction retried?
- Which product snapshot was used?
If answering those questions requires reading raw production logs manually, the system is harder to operate than it needs to be.
Build an internal transaction timeline.
Even a simple one pays for itself quickly.
Reliability comes from reducing ambiguity
The final architecture might resemble:
Client
↓
Create transaction + idempotency key
↓
Payment
↓
Queue
↓
Provider submission
↓
Processing state
↓
Webhook / reconciliation
↓
Terminal status
↓
Audit history
The goal is not to eliminate every failure.
Distributed systems do not offer that luxury.
The goal is to make every failure land in a state the system understands and can safely recover from.
That is a much stronger guarantee than hoping the external API always returns 200 OK.
AI disclosure: This article was prepared with AI assistance. The publishing editor should review the implementation examples and factual accuracy before publication.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.