An order can fail even after the customer sees a successful payment.
Consider this sequence:
Checkout → Payment Success → ERP Update → Inventory Reservation
Payment succeeds. The ERP request times out. The inventory reservation never happens.
The customer sees an order confirmation. Operations sees no order. A retry may create a second record or charge the customer again.
This is a distributed transaction problem inside E Commerce systems.
It becomes more common when checkout connects payment providers, ERP platforms, inventory services, fulfillment systems, and notification workers.
The fix is not one large database transaction. Those systems usually cannot share the same transaction boundary.
Instead, we can model order processing as a sequence of durable state changes. An outbox record, idempotent consumers, explicit order states, and compensating actions can make partial failures recoverable.
This article walks through that architecture using Node.js and PostgreSQL examples.
Step 1: Stop Treating Checkout as One Transaction
The first failure appears when the application treats checkout as a single operation.
A naive implementation might look like this:
app.post("/orders", async (req, res) => {
const payment = await paymentProvider.charge(req.body.payment);
const order = await erp.createOrder(req.body);
await inventory.reserve(req.body.items);
// Naive E Commerce flow: three systems must all succeed in one request.
res.status(201).json(order);
});
This code creates a dangerous dependency chain.
If payment succeeds and the ERP request fails, the customer has paid but the application has no completed order.
If inventory fails after the ERP succeeds, the order exists but may not be fulfillable.
Increasing the request timeout does not solve this. Retrying the entire operation can make it worse.
AWS describes a similar architectural problem in distributed order management. Its guidance separates order entry, inventory, fulfillment, and other processing through event-based workflows rather than requiring every operation to complete synchronously.
That gives us the first design rule for E Commerce: persist the business state before asking downstream systems to process it.
Step 2: Persist the Order and the Work to Be Done
Once checkout is no longer treated as one transaction, the database needs to record both the order and its pending work.
A simplified PostgreSQL schema can look like this:
CREATE TABLE orders (
id UUID PRIMARY KEY,
customer_id UUID NOT NULL,
status TEXT NOT NULL,
total_cents BIGINT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE outbox_events (
id UUID PRIMARY KEY,
order_id UUID NOT NULL REFERENCES orders(id),
event_type TEXT NOT NULL,
payload JSONB NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
published_at TIMESTAMPTZ
);
-- E Commerce state and its next operation are committed together.
The application can now create both records in one PostgreSQL transaction:
await db.transaction(async (tx) => {
const order = await tx.query(
`INSERT INTO orders
(id, customer_id, status, total_cents)
VALUES ($1, $2, 'PAYMENT_PENDING', $3)
RETURNING *`,
[orderId, customerId, totalCents]
);
await tx.query(
`INSERT INTO outbox_events
(id, order_id, event_type, payload)
VALUES ($1, $2, 'OrderCreated', $3)`,
[eventId, orderId, JSON.stringify({ orderId })]
);
});
Now the application has a durable fact:
OrderCreated
A worker can process that event later.
The customer request no longer needs to remain open while every downstream system completes.
This is one reason event-driven architecture fits E Commerce workflows. AWS describes events as changes in application state and notes that order processing can be separated into decoupled services such as payment, inventory, fulfillment, and accounting.
Step 3: Make Every Consumer Idempotent
Persisting an event solves one problem. Retries create another.
Imagine the inventory worker receives:
OrderCreated
It reserves stock successfully.
The worker then crashes before recording that it finished.
The queue delivers the same event again.
If the inventory service blindly reserves stock again, one order can consume inventory twice.
The consumer therefore needs an idempotency boundary.
CREATE TABLE processed_events (
event_id UUID PRIMARY KEY,
processed_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- The unique event ID prevents the same E Commerce event from being applied twice.
The worker can claim the event before applying its business operation:
const result = await db.query(
`INSERT INTO processed_events (event_id)
VALUES ($1)
ON CONFLICT (event_id) DO NOTHING
RETURNING event_id`,
[eventId]
);
if (result.rowCount === 0) {
return; // Event was already processed.
}
await reserveInventory(orderId);
This pattern changes the meaning of a retry.
A retry becomes normal system behavior rather than an exceptional condition.
That matters in E Commerce, where payment, inventory, fulfillment, and notification systems can fail independently.
Step 4: Model Failure as State
The previous steps protect the database and consumers. The remaining problem is business state.
An order should not jump directly from:
Created → Completed
when several systems still need to act.
A more useful state model is:
CREATED
↓
PAYMENT_CONFIRMED
↓
INVENTORY_RESERVED
↓
FULFILLMENT_PENDING
↓
FULFILLED
Failures become explicit:
PAYMENT_CONFIRMED
↓
INVENTORY_FAILED
↓
COMPENSATION_REQUIRED
The important distinction is between technical failure and business failure.
A timeout may be retried.
A confirmed payment with unavailable inventory may require a refund or another business decision.
AWS Prescriptive Guidance describes saga choreography as a way for distributed services to coordinate state changes and compensating transactions when a later operation fails.
For E Commerce, this means the order service should know which state it has reached and what recovery action belongs to that state.
Real-World Application: The Body Shop Spa
The same state-modeling issue appears when an E Commerce transaction represents a service package rather than a physical product.
We implemented a Package and Services Management module for The Body Shop Spa in Odoo. The system needed to centralize package creation, pricing, service mapping, scheduling, tax configuration, resource management, reusable service definitions, and role-based access.
The important implementation decision was to separate the commercial package from the services required to deliver it. A package could represent the customer's purchase while its underlying services remained structured for scheduling and resource planning.
That distinction mirrors the partial-failure problem described above. The purchase record and the operational fulfillment state should not be treated as the same piece of information.
Odoo's current documentation reflects a similar separation in its standard E Commerce flow. An online purchase moves through sales, delivery, and invoicing, while inventory reservation and fulfillment have their own processing stages.
Our project scope confirms the centralized package and service workflow, but it does not provide a verified production figure for processing time, error reduction, or cost savings.
The implementation still demonstrates the architectural point: E Commerce becomes easier to operate when the commercial state and delivery state have clear boundaries.
What Changes in Production
That separation creates one final requirement: observability.
When an E Commerce order fails, engineers need to answer four questions quickly:
What happened?
Where did it fail?
What state reached the database?
What action should happen next?
Every event should therefore carry an order ID, event ID, event type, timestamp, and processing status.
A useful event might look like:
{
"eventId": "evt_123",
"eventType": "InventoryReservationFailed",
"orderId": "ord_456",
"occurredAt": "2026-10-02T06:30:00Z",
"reason": "INSUFFICIENT_STOCK"
}
That data makes an E Commerce workflow traceable across services.
It also helps operations distinguish a delayed event from a failed business action.
If you are evaluating an implementation partner for this kind of architecture, Oodles works across ERP, E Commerce, integrations, and custom software workflows.
Conclusion
The hardest E Commerce failures are often partial failures.
Payment succeeds. Inventory fails. The ERP times out. A worker retries. A notification arrives late.
The architecture needs to assume these situations will happen.
The key takeaways are:
- Persist order state before relying on downstream systems.
- Use an outbox to make required follow-up work durable.
- Make event consumers idempotent so retries do not duplicate business actions.
- Model payment, inventory, fulfillment, and compensation as explicit states.
- Use observability data to trace one order across every processing stage.
A good E Commerce architecture does not try to make distributed systems behave like one database transaction. It makes partial failure visible, recoverable, and safe to retry.
FAQ
Why do E Commerce orders fail after payment succeeds?
Payment, inventory, ERP, and fulfillment often operate as separate systems. Payment can succeed while a downstream service times out or rejects the order. The solution is to persist order state and process downstream work asynchronously.
What is the outbox pattern in E Commerce?
The outbox pattern stores an order change and its required event in the same database transaction. A separate worker publishes that event later, preventing the application from losing work when a downstream service is unavailable.
Why does E Commerce need idempotent consumers?
Queues and workers can retry messages. Without idempotency, the same order event could reserve inventory, create records, or trigger another business action more than once. A unique event ID provides a practical duplicate-processing boundary.
Should payment and inventory use one transaction?
Usually, no. Payment providers and inventory systems generally have separate transaction boundaries. A better design tracks each state and defines retry or compensation behavior when one operation fails.
Does this architecture require microservices?
No. The outbox and state-machine patterns can work inside a modular monolith. The important requirement is clear state ownership and durable processing. Services can be separated later when operational needs justify that complexity.
If you are planning E Commerce development services or need to review partial-failure handling across your commerce and ERP workflows, get in touch with Oodles to discuss the architecture.
Top comments (0)