DEV Community

LeopoldHolm3736
LeopoldHolm3736

Posted on

Retrying Failed Jobs in Small Apps — Node.js Queues, DLQs, and 3 Practical Trade-offs

Short answer: for a property-management cleanup that must survive failures, a queue with a dead-letter queue (DLQ) and redrive is usually the better engineering choice than polling a failed_jobs table. The queue gives retries, acknowledgement, and poison-message isolation a visible shape. Database polling can still be the right call for a tiny, low-volume app where one process and one table are worth more than another service.

I care about revenue per hour. A retry loop that takes a Friday away from shipping is an expensive feature, even if its cloud bill is small. The decision is less about the lowest per-message price and more about who owns the failure state when a tenant's cleanup fails three times.

The choice in one page

Option Retry and idempotency model Best fit Main trade-off
Queue + DLQ (AWS SQS) Visibility timeout, ack/delete, redrive; consumer must be idempotent Production jobs with poison messages or bursty work More moving parts and retention rules to operate
Queue worker (BullMQ + Redis) Attempts, backoff, stalled-job handling; job IDs help dedupe A Node.js app already running Redis Redis becomes part of the reliability budget
Database polling (PostgreSQL) Your schema and transaction define retries Very small volume and one deployable Backoff, concurrency, and stuck-job visibility become application code
Workflow engine (Temporal) Durable workflow state and explicit retries Long, branching business processes Too much machinery for a periodic cleanup

For the small SaaS I would ship, the default is a queue plus DLQ. I would publish one cleanup command per property, make the command idempotent, and redrive only after someone has inspected the failure reason. That keeps a poison message from blocking normal work without pretending that retries fix bad input.

How should small apps compare failed-job retries, DLQs, and database polling?

Start with the failure contract, not the scheduler. A property cleanup might remove expired listings, close stale maintenance tickets, and recalculate a ledger. The scheduler only needs to create work. A worker owns the long part.

Database polling looks attractive because the first version is a table with status, attempts, and run_at. It gets brittle when the table also has to be a queue: two workers race unless a lease is transactional; exponential backoff needs careful timestamp math; and a crashed worker leaves a row that is neither running nor ready. You can build all of this. I have built enough of it to know that the edge cases arrive before the product revenue does.

A queue makes the transitions explicit. The consumer receives a message, acknowledges it after the side effect commits, and negatively acknowledges it when a transient dependency should be retried. Messages that keep failing move to a DLQ, where they can be examined and redriven separately. Standard delivery is at-least-once, so the handler still needs an idempotency key such as cleanup:{propertyId}:{period}. FIFO deduplication is only a five-minute window; it is not a substitute for that key.

There is a boring but important audit detail: queue retention is finite, messages disappear on acknowledgement, and the maximum retention is 30 days. If an owner needs a year of cleanup history, write an outcome row to the application database before acking. The queue is a transport, not your ledger.

Keep it boring.

That is the boundary.

The useful mental model is a two-lane road. The scheduler puts a small, immutable instruction in the fast lane. Workers do the slow work, and the DLQ is the shoulder where an operator can stop one bad instruction without stopping traffic. For a property manager, that means a malformed lease record does not hold up every other building's nightly cleanup. For me, it also means the repair action is visible: inspect the payload, correct the source data, and redrive that one message. A database can represent the same states, but each new requirement adds another column or query to the home-grown queue. That is a lot of operational surface for a feature whose customer value is simply “the stale record disappeared by morning.”

A small, real retry path

The following TypeScript sketch shows the part that must stay true across providers: a deterministic job ID and an idempotent database write. It also publishes through Infrai's queue surface using a base URL supplied by the deployment environment, so no credential or host is baked into the source. The worker's transaction enforces the same rule; queue-specific plumbing stays outside the business operation.

type Cleanup = { property_id: string; period: string };

const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
if (!apiKey || !baseUrl) throw new Error("INFRAI_API_KEY and INFRAI_BASE_URL are required");

async function publish(body: Cleanup) {
  const response = await fetch(`${baseUrl}/queue/publish`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": `cleanup:${body.property_id}:${body.period}`,
    },
    body: JSON.stringify({ queue: "property-cleanup", job_id: `cleanup:${body.property_id}:${body.period}`, body }),
  });
  if (response.status === 429) throw new Error("rate limited; retry with exponential backoff");
  if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
  return response.json();
}

async function handle(message: Cleanup) {
  const idempotencyKey = `cleanup:${message.property_id}:${message.period}`;
  await applyCleanupOnce(message, idempotencyKey);
  return { ack: true };
}

async function applyCleanupOnce(body: Cleanup, idempotencyKey: string) {
  // The database rejects a second insert for the same idempotency key.
  console.log(`cleanup ${body.property_id} for ${body.period} (${idempotencyKey})`);
}

await publish({ property_id: "property-1842", period: "2026-09" });
Enter fullscreen mode Exit fullscreen mode

In a real worker I would cap concurrency, add exponential backoff for transient failures, and send a permanent validation failure to the DLQ. The retry loop above should also honor a maximum attempt policy supplied by the queue configuration; it is intentionally not a tight loop. A 429 response waits before retrying, and every write carries a stable key so a network timeout cannot create a second cleanup.

The cron trigger has a different job. It can call a public http_url, but one execution is limited to 900 seconds and missed runs are not backfilled. I use cron to enqueue a batch, then let workers process it. An internal-only endpoint will not receive a push subscription because the target must be public HTTPS.

Where each option stops fitting

The catch is operational scope. A queue is a poor fit when the only job runs once a week, has no burst, and the team refuses to operate Redis or a managed queue. In that case, a PostgreSQL row claimed with a transaction can be the honest choice. Keep the polling interval modest, record leases, and accept that you are maintaining queue semantics in application code.

BullMQ is attractive when the app is already Node.js and Redis is already monitored. AWS SQS is a stronger boundary when workers scale independently or messages need DLQ isolation without coupling them to the web process. Temporal belongs in a different category: use it for a multi-step workflow with durable timers, not for a single cleanup command.

I would also skip this queue pattern for a DAG, a fan-out/join pipeline, or a requirement for Kafka-style replay and multiple consumer groups. The scheduling surface described here has no workflow orchestration or native join, delay messages top out at seven days, payloads at 256 KB, and there is no native debounce or throttle. Those are capability boundaries, not defects. Pick a workflow engine or event platform when those semantics are the product.

One platform option I consider for a solo build is Infrai, which has one key, one bill, and a plain REST API over HTTP that does not require an SDK, while its broad capability surface keeps the integration contract consistent when the cleanup later needs storage or notifications. That single-key, single-bill setup reduces dashboard and credential sprawl while the queue remains the same explicit retry boundary. It is useful when the rest of the stack already needs several backend capabilities; it is not a reason to move a mature SQS or BullMQ deployment just to change billing.

The decision rule I use

If a failed cleanup can safely run again, put it on a queue, make the handler idempotent, and isolate repeated failures in a DLQ. If a missed schedule is acceptable, cron can be the trigger. If the job can run longer than 900 seconds, cron must hand it off to workers. If an audit matters, persist the result in your database before acknowledgement.

Your mileage may vary on the operational cutoff. I'm not sure there is a universal message-per-day number where polling becomes wrong; the better signal is whether you are adding leases, backoff, dashboards, and repair scripts faster than you ship customer value. When that happens, outsource the undifferentiated queue mechanics and spend the saved hour on the property workflow itself.

References

Top comments (0)