Patterns for payment systems that stay correct at ten orders a day and at ten thousand an hour
Most payment integrations are built for the first hundred orders. They work fine until a big sale, a viral campaign or a festival weekend multiplies traffic by fifty.
Then the weaknesses show up. Webhooks time out. Duplicate events create duplicate shipments. A slow bank response blocks a worker thread. A retry creates a second charge.
This article covers the design patterns that keep payment integrations correct under load, with a focus on three areas: API calls, webhooks and failure handling.
Principle 1: Correctness first, then speed
In most systems, a small error rate is acceptable. In payments, it is not. A customer charged twice, or an order shipped without payment, costs trust and money.
So the goal is not just to handle more traffic. It is to handle more traffic while guaranteeing that every payment is processed exactly once in effect, even if individual messages are delivered more than once.
The tools for that are idempotency, state machines and asynchronous processing.
Designing outbound API calls
Your system calls the payment provider's API to create orders, capture payments, issue refunds and fetch statuses. At scale, these calls need care.
Set explicit timeouts
Never rely on default timeouts, which can be very long. Set connection and read timeouts that suit each operation, so one slow response cannot tie up your workers.
Retry carefully
Retry on network errors and server-side errors with exponential backoff and jitter. Do not retry blindly on client errors like validation failures, since the result will not change.
Make retries safe
A timeout does not mean the request failed. It means you do not know. Before retrying an operation that creates something, such as an order or a refund, check whether the first attempt succeeded. Use your own unique reference, such as an internal order ID or refund ID, so you can look it up and avoid duplicates.
Respect rate limits
Payment APIs enforce rate limits. Spread bulk operations like reconciliation or mass refunds over time, and back off when you receive rate limit responses.
Isolate failures
Use circuit breakers around payment API calls. If the provider is degraded, fail fast and queue work for later rather than stacking up requests that will time out.
Designing inbound webhooks
Webhooks are how your provider tells you what happened. At scale, they arrive in bursts, sometimes out of order and sometimes more than once. Razorpay's webhook best practices describe these behaviours clearly, and they are typical across providers.
Acknowledge fast, process asynchronously
Your webhook endpoint should do three things: verify the signature, persist the raw event, and return a 2xx response. Everything else belongs in a background worker.
This matters because providers treat slow responses as failures. Razorpay, for instance, expects a response within 5 seconds and retries with exponential backoff for up to 24 hours. If deliveries keep failing for that long, the webhook is disabled. A slow endpoint during a traffic spike can therefore create a storm of retries that makes things worse.
Deduplicate by event ID
Store each event's unique ID when you persist it, with a unique constraint in your database. If an insert fails because the ID already exists, you have already received that event and can safely ignore it.
Handle ordering with a state machine
Do not assume events arrive in the order they happened. Model each order and payment as a state machine with allowed transitions. When an event arrives, apply it only if it moves the entity forward. A late authorised event arriving after a captured event should be recorded but should not change the state.
Partition work by entity
When processing events in parallel, route events for the same order or payment to the same worker or partition. This avoids race conditions where two workers update the same order at once.
Failure handling patterns
The outbox pattern
When a payment is confirmed, you often need to update your database and trigger downstream actions like sending emails, updating inventory or notifying a warehouse. If these happen separately, a crash between them leaves your system inconsistent.
The outbox pattern solves this. Write the state change and a record of the downstream message in the same database transaction. A separate process reads the outbox and publishes messages, retrying until each is delivered.
Reconciliation as a safety net
Even well-designed systems miss events occasionally. Run periodic reconciliation jobs that compare your records with the provider's. Look for orders stuck in pending, payments captured without a matching paid order, and refunds initiated but not confirmed. Fix discrepancies automatically where possible and alert on the rest.
Mind the capture window
If you use manual capture, for example to confirm inventory before charging, remember that authorised payments are not held forever. Razorpay's capture settings documentation notes that authorised payments must be captured within a set window, and uncaptured payments are refunded automatically. A backlog in your capture queue during peak traffic can therefore turn into mass auto-refunds. Monitor capture lag closely.
Dead letter queues
Some events will fail processing repeatedly because of bugs or unexpected data. Rather than retrying forever, move them to a dead letter queue after a set number of attempts, alert the team, and reprocess once fixed.
Graceful degradation
Decide in advance what happens when parts of the system fail. If your email service is down, payments should still be confirmed. If the payment provider is degraded, your storefront should show a clear message rather than hanging.
Observability
You cannot fix what you cannot see. At minimum, track:
- Payment success rate by method and bank
- API latency and error rates by endpoint
- Webhook receipt rate, processing lag and failure count
- Number of orders in each state, especially pending
- Reconciliation mismatches
- Dead letter queue size
Use correlation IDs, such as your internal order ID, across logs so you can trace a single payment through every system it touches.
Preparing for peak events
- Load test your checkout and webhook endpoints at several times expected peak
- Scale webhook receivers and workers ahead of time
- Freeze non-essential deployments during sales
- Keep runbooks ready for common incidents like provider degradation or webhook backlog
- Coordinate with your payment provider if you expect unusually high volume
Final thoughts
Scalable payment integrations are built on a small set of ideas: every operation is idempotent, every entity has a clear state machine, slow work happens asynchronously, and reconciliation catches whatever slips through.
None of these is complicated on its own. Applied consistently, they give you a system that stays correct when traffic spikes, networks fail and events arrive in the wrong order.
Top comments (2)
Webhook retries without a signed event log are just hope with backoff.
If a settlement fails at hop 3, can a merchant reconstruct the path with queryable receipts, or only with a dashboard screenshot that dies on redesign?
Green tick is not an audit trail. marker2328h
Some comments may only be visible to logged-in visitors. Sign in to view all comments.