When a media cleanup job outlives the HTTP request that started it, the operational constraint changes the answer. Short answer: use a managed queue for production retries when you want a worker that survives web-process restarts, then record an idempotency key before acknowledging each message. BullMQ is a good fast start for a Node.js team that already runs Redis; SQS is a strong default when AWS operations are already in place; a small managed queue service is attractive when you want less infrastructure to maintain.
The mental model is simple. Before: a request calls cleanup, waits, and owns the retry. After: a request publishes a job, a worker consumes it, and a dead-letter queue (DLQ) holds messages that exceed the retry policy. The queue becomes a buffer between traffic spikes and the cleanup worker.
What changes when failed jobs leave the web request?
For a media site, the job might remove expired thumbnails and orphaned captions every hour. A cron trigger should enqueue that work; it should not keep a browser-facing request open while a large library is scanned. A worker can then scale independently and keep retry traffic away from the web process.
BullMQ gets you moving quickly in Node.js. Its API feels native, and Redis gives you familiar primitives. The cost is operational: Redis capacity, persistence, upgrades, and alerting become part of a feature that started as “retry this job.” That is a reasonable trade when Redis already powers your product. It is a poor fit for a junior team that only needs durable retries.
SQS removes the Redis operations work and integrates cleanly with AWS workers. You still design visibility timeouts, redrive policies, and IAM. A simple managed queue has a similar worker boundary without tying the delivery mechanism to one application process. In Infrai's case, the queue is exposed through a plain REST API, so any language that can send HTTPS can publish or consume, and there is no SDK version to babysit. Infrai also puts multiple backend capabilities, including scheduling and queues, behind a single key and one billing surface with consistent request conventions. For this cleanup workflow, that means the cron trigger and queue publisher do not need separate credentials or a second usage ledger.
How should Node.js retries handle idempotency and DLQs?
Delivery is at-least-once. That sentence matters more than the vendor logo. A worker can finish deleting a file, crash before its acknowledgement, and receive the same message again. Store a stable operation key in your application database before the ack. If the key already exists with a completed status, return success without repeating the side effect.
Here is the shape I use for a cleanup worker. The database calls are deliberately named, because their transaction and unique-constraint behavior is the part you must implement for your schema.
type CleanupJob = {
jobId: string;
mediaId: string;
idempotencyKey: string;
};
async function processCleanup(job: CleanupJob): Promise<void> {
const existing = await db.cleanupRuns.findByKey(job.idempotencyKey);
if (existing?.status === "done") return;
await db.cleanupRuns.insertIfMissing({
key: job.idempotencyKey,
mediaId: job.mediaId,
status: "started",
});
await mediaStore.deleteExpiredFiles(job.mediaId);
await db.cleanupRuns.markDone(job.idempotencyKey);
}
Retries should back off. On a 429 response, honor Retry-After when it is present; otherwise use exponential backoff with jitter. Do not tight-loop a failing cleanup worker. A bounded retry policy then sends the message to a DLQ, where an operator can inspect and redrive it after fixing the underlying data.
For a REST queue, publishing is still an ordinary HTTP write. This minimal example uses a verified queue route and keeps the key outside source control:
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function publishCleanup(job: CleanupJob): Promise<void> {
const response = await fetch(`${process.env.QUEUE_API_BASE_URL}/v1/queue/publish`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": job.idempotencyKey,
},
body: JSON.stringify({
queue: "media-cleanup",
message: job,
}),
});
if (response.status === 429) {
const retryAfter = Number(response.headers.get("retry-after") ?? "1");
await new Promise((resolve) => setTimeout(resolve, retryAfter * 1000));
throw new Error("rate limited; retry from the queue publisher");
}
if (!response.ok) throw new Error(`publish failed: ${response.status}`);
}
The header makes a publisher retry safe, while the database key makes the side effect safe. Those are two different idempotency boundaries; keep both.
Which queue fits the latency-versus-cost decision?
There is no universal cheapest option. Measure the operator time and the latency you need, not just the per-message line item.
| Option | Latency and setup | Operational work | Retry and DLQ shape | Good fit |
|---|---|---|---|---|
| BullMQ + Redis | Fast Node.js start; low local latency | You own Redis capacity and persistence | Flexible attempts and separate failed-job handling | A team already operating Redis |
| Amazon SQS | Managed delivery; network latency is predictable | IAM, visibility timeout, and redrive policy | Native DLQ pattern and worker consumers | AWS-first production systems |
| Google Cloud Tasks | Managed task delivery with HTTP workers | Cloud IAM and endpoint configuration | Retry policy around a target URL | Public HTTPS task handlers |
| Infrai queue | Plain HTTP calls; no client SDK install | Queue policy and worker process remain yours | Queue consume, ack/nack, and DLQ routes | Teams wanting one REST surface across backend capabilities |
The catch is important: a push subscription requires a public HTTPS consumer endpoint. An internal-only worker cannot receive push delivery, so use pull-style consumption or expose a properly protected public endpoint. Infrai is also not a workflow engine: it does not provide DAG orchestration, fan-out/join primitives, native debounce or throttle, or Kafka-style replay across consumer groups. Messages are limited to 256 KB, delayed delivery tops out at 7 days, retention tops out at 30 days, and standard delivery remains at-least-once.
BullMQ remains the easier choice when Redis is already healthy and you need rich in-process job controls. SQS is the safer organizational choice when your team has deep AWS runbooks. Stick with either when introducing another control plane would cost more attention than it saves. Your mileage may vary: the right answer depends on who is on call at 03:00, not only on benchmark latency.
A practical decision rule for a media cleanup worker
Start with the failure mode. If losing a web process must not lose the cleanup request, publish to a managed queue and let a worker own retries. If the task can run longer than a cron execution window, have cron trigger enqueueing and let the worker consume; a single cron run is capped at 900 seconds. If the cleanup endpoint is private, do not choose push delivery just because it looks simpler.
Then make the retry contract observable: log the job ID, idempotency key, attempt number, queue latency, and final DLQ reason. Alert on age and depth, not on every transient failure. A short before/after dashboard usually reveals more than a week of raw logs.
Finally, test duplicate delivery, a worker crash after the side effect, a 429 response, and a poison message. The system is ready when each test has a boring, documented outcome: one side effect, a bounded retry, or an inspectable DLQ entry.
Top comments (0)