DEV Community

Cover image for MobileTopUP: Designing a Reliable Recharge Transaction Workflow
MobileTopUP
MobileTopUP

Posted on

MobileTopUP: Designing a Reliable Recharge Transaction Workflow

Distributed transactions become interesting exactly when the happy path stops being reliable.

A basic recharge implementation might look like this:

charge card
    ↓
call recharge provider
    ↓
return success
Enter fullscreen mode Exit fullscreen mode

In a local test environment, that can appear perfectly adequate.

Production introduces:

  • duplicate requests;
  • timeouts;
  • delayed provider responses;
  • asynchronous status updates;
  • partial failures;
  • process crashes;
  • payment success followed by fulfilment uncertainty;
  • users refreshing the page;
  • workers retrying jobs.

The underlying prepaid mobile top-up use case is a useful example because a single user action can involve payment, external fulfilment, asynchronous confirmation and an operation that should never be executed twice accidentally.

The solution is not to add more try/catch blocks.

The transaction needs an explicit lifecycle.

Start with a transaction record before external execution

Do not wait until the provider responds successfully before creating a database record.

Create the transaction first.

For example:

{
  "id": "txn_123",
  "status": "created",
  "recipient": "+447700900123",
  "quote_id": "q_456",
  "idempotency_key": "abc-123",
  "created_at": "..."
}
Enter fullscreen mode Exit fullscreen mode

Now every later operation has an internal identity.

That identity can be used in:

  • logs;
  • payment metadata;
  • provider metadata;
  • support tools;
  • reconciliation jobs.

A transaction that fails midway still exists.

That is valuable.

Use a state machine instead of a success flag

A boolean cannot represent asynchronous processing well.

This:

success = null
Enter fullscreen mode Exit fullscreen mode

quickly becomes ambiguous.

A state model is clearer.

For example:

created
payment_pending
payment_authorized
submitted
processing
succeeded
failed
Enter fullscreen mode Exit fullscreen mode

Depending on the architecture, you might also need:

payment_failed
cancelled
refunded
manual_review
Enter fullscreen mode Exit fullscreen mode

The exact state names matter less than defining what each one means.

For every state, ask:

  • What must already have happened?
  • Which transitions are allowed?
  • Is the state terminal?
  • Can a worker safely retry from here?

Idempotency belongs at the transaction boundary

Imagine this request:

POST /api/recharges
Enter fullscreen mode Exit fullscreen mode

The user double-clicks.

Or the frontend retries because the first response takes too long.

Without protection:

Request 1 → recharge
Request 2 → recharge
Enter fullscreen mode Exit fullscreen mode

The customer may pay twice or the recipient may receive duplicate value.

Instead, require an idempotency key:

Idempotency-Key: 70356c3a-...
Enter fullscreen mode Exit fullscreen mode

Store it with a unique database constraint.

Pseudo-code:

existing = find_transaction(idempotency_key)

if existing:
    return existing

transaction = create_transaction(idempotency_key)
process(transaction)
Enter fullscreen mode Exit fullscreen mode

The unique constraint matters because two concurrent application workers can otherwise both pass an application-level “does this exist?” check.

Correctness should survive concurrency.

Do not blindly retry an ambiguous provider request

Suppose you send:

POST /provider/recharge
Enter fullscreen mode Exit fullscreen mode

The provider completes the recharge.

Your connection drops before the HTTP response arrives.

From your application's perspective:

request timeout
Enter fullscreen mode Exit fullscreen mode

From the provider's perspective:

recharge succeeded
Enter fullscreen mode Exit fullscreen mode

If your retry policy says:

timeout → retry immediately
Enter fullscreen mode Exit fullscreen mode

you may create a duplicate.

A timeout means:

We do not know the outcome.

That is different from:

The operation failed.

If the provider supports idempotency, use it.

If it provides transaction lookup, query the original request.

If it sends asynchronous callbacks, wait for the callback within a reasonable reconciliation window.

Retries need business context.

Give every provider operation a correlation identifier

A useful integration keeps both sides traceable.

For example:

internal transaction:
txn_123

provider request reference:
mobiletopup-txn_123

provider transaction:
ext_987
Enter fullscreen mode Exit fullscreen mode

Store the provider transaction ID as soon as you receive it.

Then support and reconciliation processes can move between:

your system ↔ provider system
Enter fullscreen mode Exit fullscreen mode

without searching by amount and timestamp.

Correlation IDs are cheap.

Debugging without them is not.

Separate payment state from recharge state

One of the most important modeling decisions is not to treat payment success as transaction success.

Consider:

Payment: authorized
Recharge: processing
Enter fullscreen mode Exit fullscreen mode

That is a perfectly valid intermediate state.

Likewise:

Payment: failed
Recharge: not submitted
Enter fullscreen mode Exit fullscreen mode

is very different from:

Payment: captured
Recharge: provider rejected
Enter fullscreen mode Exit fullscreen mode

A transaction may therefore contain separate dimensions:

{
  "payment_status": "captured",
  "fulfilment_status": "processing"
}
Enter fullscreen mode Exit fullscreen mode

The overall user-facing status can be derived from those states.

Do not destroy useful information by compressing both into one column too early.

Make asynchronous completion normal

External fulfilment often works better as an asynchronous workflow.

A simplified architecture:

API request
   ↓
create transaction
   ↓
authorize/capture payment
   ↓
enqueue recharge job
   ↓
provider submission
   ↓
processing
   ↓
webhook or polling
   ↓
final state
Enter fullscreen mode Exit fullscreen mode

The API does not need to keep an HTTP connection open for the entire provider lifecycle.

The frontend can poll:

GET /api/recharges/txn_123
Enter fullscreen mode Exit fullscreen mode

or receive updates through another mechanism.

This also makes temporary provider slowness easier to absorb.

Webhooks need the same defensive engineering

A webhook can arrive:

  • once;
  • twice;
  • out of order;
  • much later than expected.

Treat webhook processing as idempotent.

For example:

provider event ID
+
unique constraint
Enter fullscreen mode Exit fullscreen mode

If the same event arrives twice, processing it twice should not produce duplicate side effects.

Also validate the webhook's authenticity using the provider's supported verification mechanism.

Do not trust an arbitrary public POST request simply because it contains a transaction ID.

Add reconciliation even if webhooks exist

Webhooks are useful.

They are not magic.

Events can be delayed or lost because of:

  • endpoint downtime;
  • configuration mistakes;
  • provider incidents;
  • networking problems;
  • internal processing bugs.

A reconciliation worker can periodically find transactions such as:

status = processing
AND updated_at < now - threshold
Enter fullscreen mode Exit fullscreen mode

and query the provider.

This gives the system a recovery path that does not depend on every asynchronous event arriving perfectly.

Store events, not only the latest state

A row that currently says:

status = succeeded
Enter fullscreen mode Exit fullscreen mode

does not explain how it got there.

An event history might say:

12:00:01 transaction_created
12:00:05 payment_authorized
12:00:06 provider_submitted
12:00:07 provider_acknowledged
12:00:24 provider_processing
12:00:32 recharge_succeeded
Enter fullscreen mode Exit fullscreen mode

That timeline is enormously useful.

You can implement it with:

  • a dedicated transaction-events table;
  • structured application events;
  • or an appropriate event-sourcing approach where warranted.

You do not need full event sourcing merely to maintain an audit trail.

Retries should be state-aware

A retry policy such as:

retry every exception three times
Enter fullscreen mode Exit fullscreen mode

is dangerous for financial or fulfilment operations.

Instead, decide by operation type.

A read-only provider lookup?

Usually safe to retry.

Submitting a recharge with provider idempotency?

Potentially safe to retry using the same key.

Submitting without provider idempotency after an ambiguous timeout?

Reconcile first.

Retry logic belongs to the business workflow, not just the HTTP client.

Design for support from day one

Operational support will eventually need answers to questions such as:

  • Did the payment complete?
  • Was the recharge submitted?
  • What external ID did the provider return?
  • When was the last provider status check?
  • Did a webhook arrive?
  • Was the transaction retried?
  • Which product snapshot was used?

If answering those questions requires reading raw production logs manually, the system is harder to operate than it needs to be.

Build an internal transaction timeline.

Even a simple one pays for itself quickly.

Reliability comes from reducing ambiguity

The final architecture might resemble:

Client
  ↓
Create transaction + idempotency key
  ↓
Payment
  ↓
Queue
  ↓
Provider submission
  ↓
Processing state
  ↓
Webhook / reconciliation
  ↓
Terminal status
  ↓
Audit history
Enter fullscreen mode Exit fullscreen mode

The goal is not to eliminate every failure.

Distributed systems do not offer that luxury.

The goal is to make every failure land in a state the system understands and can safely recover from.

That is a much stronger guarantee than hoping the external API always returns 200 OK.


AI disclosure: This article was prepared with AI assistance. The publishing editor should review the implementation examples and factual accuracy before publication.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.