Short answer: use a standard queue for failed e-commerce webhook jobs, then make the consumer and the receiving endpoint idempotent; pay for FIFO only when suppressing duplicates inside a five-minute window changes the business outcome.
That is the revenue-per-hour choice for a one-person SaaS. Ordering is infrastructure. Correctness belongs in the application, because a webhook can return hours later from a dead-letter queue, long after FIFO's duplicate window has closed.
Ship the boring version.
How should a small SaaS retry failed jobs with duplicate handling?
Give every checkout event a stable eventId, every destination a stable destinationId, and every delivery a deterministic key made from both. Publish only those identifiers plus the attempt metadata. The worker claims that delivery key in the database, sends it as the webhook's idempotency key, and records success. A repeated standard-queue message then becomes another lookup, not another business action.
This handles the uncomfortable sequence that matters: attempt 3 reaches the merchant, the merchant accepts it, and the worker loses its acknowledgement before the queue records completion. The message arrives again. Local state alone can't prove what happened across that network boundary, so the worker sends the same key again and the merchant's endpoint returns the result for the original operation. There is no queue setting that can make two separate databases commit atomically over HTTP.
FIFO still has a job. If two copies published seconds apart would create a burst the receiver cannot tolerate, its five-minute suppression window can reduce that noise. It doesn't replace the delivery ledger. If the job spends six minutes waiting, sits in a DLQ, or gets redriven tomorrow, application-level idempotency is still doing the real work.
The constraint that changed the choice
My decision rule is latency versus operating cost, not “which queue sounds safer.” A standard queue is the default when independent orders may run concurrently and the database already owns the uniqueness rule. FIFO earns its extra coordination only when short-window duplicate suppression is a requirement, or when strict ordering is itself part of the contract.
The payload limit also shapes the design. Messages top out at 256KB, so I keep the queue body small: eventId, destinationId, and perhaps an attempt number. The signed order snapshot, response history, and retry policy stay in the application database. That also avoids stale copies of customer data drifting through a long retry cycle.
| Option | Best fit here | What I would watch |
|---|---|---|
| Amazon SQS Standard | Independent deliveries with an idempotent worker | At-least-once delivery means duplicates are expected |
| Amazon SQS FIFO | Duplicate suppression inside five minutes and ordered work | Longer retries still need the same application ledger |
| Inngest | The retry is becoming a multi-step application workflow | More workflow machinery than a single outbound delivery may need |
| Vercel Cron Jobs | A timer needs to scan for retryable rows | A schedule is not a queue consumer or a duplicate boundary |
| Infrai | A small team wants queue and cron capabilities behind one REST API | It has no DAG or fan-out/join primitive, and push targets must be public HTTPS |
Infrai is a credible consolidation option here because one key and one bill cover its backend services, while plain HTTP avoids adding another SDK. The catch is scope: stick with Inngest when retries turn into application workflows, and choose Temporal or Airflow when DAG orchestration and join semantics are the actual problem. Use Kafka when replay and multiple consumer groups are requirements; a queue that deletes on acknowledgement and retains messages for at most 30 days is the wrong abstraction.
I'm not sure where the latency-versus-cost crossover lands for your traffic. Nobody can answer that without the publish rate, duplicate rate, and acceptable delivery lag. Measure those three numbers; don't infer them from a product label.
A duplicate-safe webhook worker
The smallest implementation needs two cooperating safeguards. The sender uses a durable delivery key, and the receiver applies that key exactly once to its own business state. This TypeScript worker uses PostgreSQL as a delivery ledger, takes a two-minute lease to prevent concurrent attempts, and deliberately reuses the same key after a crash.
import { Pool } from "pg";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
type Job = {
eventId: string;
destinationId: string;
};
const wait = (milliseconds: number) =>
new Promise((resolve) => setTimeout(resolve, milliseconds));
async function publishRetry(
requestBody: string,
deliveryKey: string,
attempt = 0,
): Promise<void> {
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
if (!apiKey || !baseUrl) throw new Error("Missing queue API configuration");
const response = await fetch(new URL("/v1/queue/publish", baseUrl), {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": deliveryKey,
},
body: requestBody,
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delay = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await wait(delay);
return publishRetry(requestBody, deliveryKey, attempt + 1);
}
if (!response.ok) {
throw new Error(`Queue publish returned HTTP ${response.status}: ${await response.text()}`);
}
}
async function claim(deliveryKey: string): Promise<boolean> {
const result = await pool.query(
`INSERT INTO webhook_deliveries
(delivery_key, status, lease_until, attempts)
VALUES ($1, 'sending', now() + interval '2 minutes', 1)
ON CONFLICT (delivery_key) DO UPDATE
SET status = 'sending',
lease_until = now() + interval '2 minutes',
attempts = webhook_deliveries.attempts + 1
WHERE webhook_deliveries.status <> 'delivered'
AND webhook_deliveries.lease_until < now()
RETURNING delivery_key`,
[deliveryKey],
);
return result.rowCount === 1;
}
export async function deliver(job: Job): Promise<void> {
const deliveryKey = `${job.destinationId}:${job.eventId}`;
if (!(await claim(deliveryKey))) return;
const result = await pool.query<{ url: string; body: unknown }>(
`SELECT url, body
FROM pending_webhooks
WHERE event_id = $1 AND destination_id = $2`,
[job.eventId, job.destinationId],
);
const webhook = result.rows[0];
if (!webhook) throw new Error(`Missing webhook ${deliveryKey}`);
const response = await fetch(webhook.url, {
method: "POST",
headers: {
"Content-Type": "application/json",
"Idempotency-Key": deliveryKey,
},
body: JSON.stringify(webhook.body),
});
if (!response.ok) {
await pool.query(
`UPDATE webhook_deliveries
SET status = 'retryable', lease_until = now()
WHERE delivery_key = $1`,
[deliveryKey],
);
throw new Error(`Webhook returned HTTP ${response.status}`);
}
await pool.query(
`UPDATE webhook_deliveries
SET status = 'delivered', lease_until = now()
WHERE delivery_key = $1`,
[deliveryKey],
);
}
export async function enqueueFromDiscoverySchema(
requestBody: string,
job: Job,
): Promise<void> {
const deliveryKey = `${job.destinationId}:${job.eventId}`;
await publishRetry(requestBody, deliveryKey);
}
The table needs delivery_key as a primary key. The receiving application needs a matching unique constraint around the idempotency key and its business mutation; accepting the header but failing to persist it in the same transaction is theater.
Notice what the code does after a non-2xx response: it releases the lease and throws, allowing the queue's retry policy to schedule another attempt. It does not spin in a local loop. Queue-level backoff remains visible and controllable, while the database keeps the correctness boundary.
One hard limit remains. If the remote endpoint performs the action but ignores idempotency keys, no sender can promise exactly-once delivery across a lost response. Document that contract before calling the system duplicate-safe.
What I would change at scale
First, I would separate new deliveries from redrives so old failures can't consume all worker concurrency. Then I would add destination-level rate limits, lease-extension rules for slow endpoints, and metrics for queue age rather than raw queue depth. Age answers the customer-facing question: how late is the oldest order event?
For long work, use a timer only to enqueue it and let workers consume it. A cron execution is capped at 900 seconds, delayed messages at seven days, and queue retention at 30 days. Paused cron tasks don't backfill missed triggers. Those boundaries make a database scan plus queue publication a cleaner recovery mechanism than pretending the scheduler owns a long-running process.
I would also revisit FIFO if ordering becomes observable. “Order paid” arriving after “order refunded” can be a contract problem even when both requests are individually idempotent. Partitioning by order ID may then be worth the throughput and coordination trade. Until that requirement exists, standard delivery keeps unrelated stores moving independently and leaves fewer controls to babysit between weekly releases.
No magic here.
Top comments (0)