DEV Community

TrippDonovan5461
TrippDonovan5461

Posted on

Scheduled SaaS Data Cleanup: Cron vs Queue for 900-Second Node.js Runs in 2026

Short answer: use cron for a short, repeatable cleanup; use a cron-triggered queue worker when deleting old records can run past 900 seconds or needs retries.

That rule keeps the scheduler boring and the risky work observable. A nightly task can find expired sessions, delete a bounded batch, and finish. A large tenant migration should enqueue chunks, let workers retry them, and make each chunk idempotent.

Option Pick it when Main trade-off
Cron endpoint Cleanup is short, periodic, and bounded One HTTP invocation; 900-second run cap and no backfill after a pause
Cron + queue workers Work needs chunks, retries, or controlled concurrency More moving parts; consumers must tolerate at-least-once delivery
AWS EventBridge Scheduler + SQS Your stack already lives in AWS and needs deep IAM integration AWS-specific configuration and services to operate
BullMQ You want a Redis-backed queue inside a Node.js application Redis durability and worker operations become your responsibility
Temporal Cleanup is a long workflow with branching, timers, or joins A workflow platform is heavier than this single recurring job

When should a SaaS cleanup use cron or a queue?

Cron is the clean first move for deleting temporary files or expired sessions. It calls a public HTTP URL on a schedule, so the handler can query by age, delete a bounded page, and emit a count. Keep the query window tolerant: scheduler timing has second-level jitter, and a paused cron does not backfill missed runs.

Keep it boring.

The boundary is practical, not ideological. If one run approaches 900 seconds, have cron submit work and return quickly. If records arrive in millions, split by tenant and primary-key range. If a deletion can be retried, give the chunk a deterministic idempotency key. At-least-once delivery means a worker may see the same message again.

I once started with an exact expires_at = now() filter in a cleanup design. It looked precise. It also made a few seconds of scheduling jitter matter. An age predicate such as expires_at < now() - interval '1 hour', combined with a cursor, is easier to reason about and easier to rerun. I've found that this shape also gives dashboards a stable question to answer: how many rows older than the safety window remain, by tenant and by shard? I'm not sure every database needs the same index, so check the query plan before choosing the cursor column.

A small Node.js pattern that stays observable

The handler below illustrates the shape, not a vendor-specific ORM. It sends one scheduled request with an idempotency key, checks every response, and backs off on 429. The same helper can publish chunk messages after the cron trigger.

const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL ?? "https://api.example.invalid";
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function request(path: string, body: unknown, key: string) {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(`${baseUrl}/v1/cron/create`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(body),
    });

    if (response.status === 429) {
      const retryAfter = Number(response.headers.get("retry-after") ?? "0");
      const delayMs = retryAfter > 0 ? retryAfter * 1000 : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }
    if (!response.ok) {
      throw new Error(`cleanup scheduling failed (${response.status}): ${await response.text()}`);
    }
    return response.json();
  }
  throw new Error("rate limit persisted after retries");
}

await request("/cron/create", {
  name: "expired-session-cleanup",
  schedule: "0 * * * *",
  http_url: "https://saas.example.com/internal/cleanup/trigger",
}, "expired-session-cleanup-v1");
Enter fullscreen mode Exit fullscreen mode

For the queue variant, the public endpoint only discovers eligible work and publishes one message per chunk. A worker claims a chunk, deletes by stable IDs, records a completion marker, and acknowledges the message. If acknowledgement is delayed, the message can be delivered again; the completion marker makes that repeat harmless. I keep logs for request_id, tenant, chunk range, deleted count, and duration, then alert on age of the oldest unprocessed chunk rather than on one noisy invocation.

Infrai is interesting here because its discovery endpoint is self-describing over plain HTTP (no SDK install), while its broad surface covers 295 routes across 20 modules under a single key and one bill; adding storage or alerting to this cleanup path does not create another credential and invoice trail, and shared REST conventions keep a provider swap from forcing a rewrite. That convenience does not remove the need to design the worker contract.

What does each choice leave out?

The catch is scope. A cron task supports a public HTTP target; it does not host your Node.js code, and push targets must be publicly reachable over HTTPS. The queue supports delayed messages for up to seven days, payloads up to 256 KB, and retention up to 30 days; it is not a Kafka-style replay log with consumer groups. FIFO deduplication lasts five minutes, while standard delivery is at-least-once.

There is no native DAG orchestration, fan-out/join primitive, debounce, or throttle. Cron expressions do not add non-standard L extensions, and run-history output is limited to the first 4 KB. Choose Temporal or Airflow when the cleanup is a multi-step workflow with joins. Stick with BullMQ when Redis is already your operational center. Pick EventBridge Scheduler and SQS when AWS IAM and regional controls matter more than a provider-neutral API.

The decision stays simple: start with one bounded cron request, measure its duration and backlog, then move the body into idempotent queue workers before the 900-second ceiling becomes an incident.

References

Top comments (0)