Processing inbound e-commerce order webhooks directly against downstream ERP systems (such as Odoo or SAP) exposes core applications to third-party API latency, rate limits, and transient network failures.
The async-erp-sync-engine microservice decouples webhook ingestion from execution, accepting payloads in under 10ms while ensuring durable background processing via BullMQ and Redis.
1. System Architecture & Queue Pipeline
The microservice separates the HTTP intake layer from worker processing loops so client webhooks are never blocked:
ERP / Shopify ┌────────────────────────────────────┐
Webhook POST ────────────▶│ Express HTTP Server │
(order-created) │ POST /api/v1/webhooks/order-created│
│ │
│ 1. Zod payload validation │
│ 2. enqueueErpJob() → BullMQ │
│ 3. HTTP 202 Accepted returned │
└─────────────────┬──────────────────┘
│ enqueue
▼
┌────────────────────────────────────┐
│ Redis 7 (BullMQ backend) │
│ │
│ ┌──────────────────────────────┐ │
│ │ erp-sync-queue (active) │ │
│ └──────────────┬───────────────┘ │
│ │ │
│ ┌──────────────▼───────────────┐ │
│ │ failed set (dead-letter) │ │
│ └──────────────────────────────┘ │
└─────────────────┬──────────────────┘
│ dequeue
▼
┌────────────────────────────────────┐
│ ERP Worker (BullMQ) │
│ │
│ • Concurrency: 5 (configurable) │
│ • Rate limit: 20 jobs / second │
│ • Retry: 5 attempts, exponential │
│ backoff (1s → 2s → 4s → 8s) │
│ • Permanent failures → failed │
│ state (UnrecoverableError) │
└─────────────────┬──────────────────┘
│ sync
▼
┌────────────────────────────────────┐
│ Odoo / SAP ERP API (external) │
└────────────────────────────────────┘
2. Key Resiliency Patterns & Design Decisions
To guarantee message delivery and protect external ERP APIs from thundering herds, the microservice implements five architectural safeguards:
-
Immediate 202 Accepted responses: Inbound webhooks are validated with Zod runtime schemas and enqueued into BullMQ within milliseconds, immediately returning
HTTP 202 Accepted. - Durable BullMQ state management: Redis 7 with Lua scripts guarantees atomic job state transitions, persistent queue storage across node restarts, and worker-level rate limiting (capped at 20 jobs/sec).
-
5-attempt exponential backoff: Jobs hitting transient errors retry with increasing delays (
1s → 2s → 4s → 8s), backed by an inner retry loop for brief network blips. - Dead-letter queue (failed set): Unrecoverable errors, or jobs that exhaust all 5 attempts, move to a dedicated failed set for inspection and manual replay.
-
Graceful shutdown: The worker listens for
SIGTERM/SIGINTso in-flight jobs finish before the process exits, enabling zero-downtime rolling container deployments.
3. Queue Implementation Snippet (TypeScript)
import { Queue } from 'bullmq';
import { z } from 'zod';
import { redisConnection } from '../config/redis';
export const OrderWebhookSchema = z.object({
order_id: z.string().min(1),
customer_ref: z.string().min(1),
amount: z.number().positive(),
currency: z.string().length(3),
});
export type OrderWebhookPayload = z.infer<typeof OrderWebhookSchema>;
export const erpQueue = new Queue('erp-sync-queue', {
connection: redisConnection,
defaultJobOptions: {
attempts: 5,
backoff: {
type: 'exponential',
delay: 1000, // retries after 1s -> 2s -> 4s -> 8s
},
removeOnComplete: true,
},
});
export async function enqueueErpJob(payload: OrderWebhookPayload) {
const jobId = `order-${payload.order_id}-${Date.now()}`;
return await erpQueue.add('sync-order', payload, { jobId });
}
Originally published at ctousman.com.
About the Author:
I'm Usman Khan, Fractional CTO & Systems Architect. I advise high-growth SaaS and e-commerce platforms on high-concurrency backend architecture, distributed queues, and system performance.
Need an architectural audit of your backend queue pipelines? Book a 30-min strategy call
Top comments (0)