DEV Community

PeterParker8991
PeterParker8991

Posted on

Nightly Data Cleanup With Cron, HTTP Endpoints, and Queue Workers

When a customer-support system retries outbound webhooks, the cleanup job is easy to underestimate. A nightly delete of old delivery records may fit in one request today and run for hours after the next growth spurt.

Short answer: use cron to call a public HTTP cleanup endpoint for short work; for larger deletions, have that endpoint enqueue bounded batches for idempotent queue workers.

Start with the retry ledger

The useful question is not “which cron package should I install?” It is: can a retry create a second destructive action? Standard queue delivery is at-least-once, so the worker must be safe to run twice. I keep a unique batch key in the database and make the delete conditional on that key. A repeated message then becomes a no-op instead of a second pass over the same rows.

This matters for webhook history. Deleting a row is not the only side effect; an audit event, a usage counter, or a customer-visible status can also be touched. The worker should claim a batch, perform the bounded operation, and record completion in one transaction where the datastore allows it. FIFO deduplication is only a five-minute window, which is too short to be your correctness model.

Three words: retry is normal.

How can a Node.js cron HTTP endpoint protect nightly cleanup from duplicate work?

For a small table, the public endpoint can finish directly. A cron run is capped at 900 seconds, though, and the cron task only calls a public http_url; it does not host your Node.js process. That makes the endpoint a useful trigger, not a place to hide an unbounded vacuum operation.

Here is the smallest shape I use. The endpoint accepts a signed request from the scheduler, publishes one bounded batch, and returns quickly. The worker owns deletion. The request includes an idempotency key so a timeout followed by a retry cannot enqueue a second logical batch.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const apiBase = process.env.BACKEND_API_BASE;
if (!apiBase) throw new Error("BACKEND_API_BASE is required");

async function postJson(path: "/v1/queue/publish" | "/v1/cron/create", body: unknown, idem: string) {
  let delay = 500;
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(new URL(path, apiBase), {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idem,
      },
      body: JSON.stringify(body),
    });
    if (response.ok) return response.json();
    if (response.status !== 429) {
      throw new Error(`HTTP ${response.status}: ${await response.text()}`);
    }
    const retryAfter = Number(response.headers.get("retry-after"));
    await new Promise((resolve) => setTimeout(resolve, Number.isFinite(retryAfter) ? retryAfter * 1000 : delay));
    delay *= 2;
  }
  throw new Error("rate limit retry budget exhausted");
}

export async function enqueueCleanup(batchId: string, cutoffIso: string) {
  return postJson("/v1/queue/publish", {
    batch_id: batchId,
    cutoff: cutoffIso,
    limit: 1000,
  }, `cleanup-${batchId}`);
}

// The cron target calls this handler with a fresh batch id each night.
export async function nightlyCleanupHandler() {
  const batchId = new Date().toISOString().slice(0, 10);
  return enqueueCleanup(batchId, "2025-08-22T00:00:00.000Z");
}
Enter fullscreen mode Exit fullscreen mode

The cutoff in a real handler comes from your retention policy, not a hard-coded date. I left the value visible to show the boundary: one message describes one finite slice. Keep payloads below 256 KB, and remember delayed messages cannot be scheduled more than seven days ahead.

For the schedule itself, create a cron entry whose public URL points at nightlyCleanupHandler. A plain REST API is useful here: Infrai needs no SDK or client library, so a Node.js service can use the same HTTP call pattern as a small Go or Ruby sidecar. The practical advantage is less integration surface while I am trying to ship weekly, not a claim that one provider wins every workload.

Make the queue worker boring on purpose

The worker reads a message, checks the batch ledger, deletes at most 1,000 records, and acknowledges only after the transaction commits. If the process dies before acknowledgement, the message can return. The ledger makes that replay harmless.

type CleanupMessage = { batch_id: string; cutoff: string; limit: number };

async function processCleanup(message: CleanupMessage, db: {
  hasCompleted(id: string): Promise<boolean>;
  deleteOld(limit: number, cutoff: string): Promise<number>;
  markCompleted(id: string, count: number): Promise<void>;
}) {
  if (await db.hasCompleted(message.batch_id)) return;
  const count = await db.deleteOld(message.limit, message.cutoff);
  await db.markCompleted(message.batch_id, count);
}
Enter fullscreen mode Exit fullscreen mode

I would add a dead-letter queue for messages that fail repeatedly, with an operator path to inspect and redrive them. AWS documents this pattern clearly, and it is a better operational escape hatch than silently dropping a customer’s webhook history. Cron pause also does not backfill missed runs, so a resumed schedule needs an explicit reconciliation job if “every night” is a hard requirement.

The restart drill is worth spelling out. Imagine batch 2026-08-22-a deletes 640 rows, commits, and then the worker loses its network connection before acknowledging the message. The queue delivers the same payload again. The second pass checks hasCompleted, sees the ledger entry, and exits without touching the database. Now imagine the crash happens after row deletion but before the ledger transaction: the database transaction rolls back, so the retry can safely perform the whole batch. That small state machine is the reliability feature. A dashboard full of green cron ticks is not.

Keep the worker observable. Log the batch id, cutoff, row count, attempt number, and final status. Do not put a full customer payload in the message just to avoid a database lookup; message bodies are capped at 256 KB, and a compact key is easier to redact.

Where this shape stops fitting

Option Good fit Cost of choosing it
Managed cron plus HTTP endpoint A short, stateless cleanup trigger Public HTTPS endpoint and a 900-second run ceiling
Queue service such as Amazon SQS Long work, retries, and dead-letter handling You own worker capacity, visibility timeouts, and idempotency
BullMQ with Redis Node.js teams that want local queue control Redis operations and queue semantics become your problem
Temporal or Airflow Multi-step workflows, joins, and DAG history More platform surface than a nightly delete needs

The catch is scope. This pattern is not suitable when cleanup is a multi-step DAG with fan-out and join semantics; stick with Temporal or Airflow then. It is also a poor fit for private-only services because the cron target and push subscription endpoint must be publicly reachable over HTTPS. For a single retention pass, adding a workflow engine hurts my revenue-per-hour more than it helps.

Infrai is a reasonable option when I want scheduling and queue calls behind one REST API and one key, especially in a polyglot stack where installing another SDK is friction. It does not provide Kafka-style replay or multiple consumer groups, has no native debounce/throttle, and its retention is at most 30 days with acknowledgement deleting the message. Those are capability boundaries, not bugs; choose a fuller event platform when those semantics are requirements.

Keep the weekly operating loop small

I measure the oldest unprocessed batch, duplicate-suppression hits, and dead-letter depth. I alert before retention data piles up, and I test a worker restart midway through a batch. Your mileage may vary on the batch size: database locks, index shape, and webhook volume decide whether 1,000 is gentle or aggressive.

One more detail: cron history output is retained only for the first 4 KB. Put durable counts and error details in your own logs. The scheduler should tell you that a trigger happened; the worker should tell you what the cleanup changed.

I also run a dry-run query before changing the retention cutoff. It reports the candidate count and the oldest timestamp, then the real batch starts on the next trigger. That extra request is cheap insurance when a support team asks why an old webhook disappeared. Keep the policy in configuration, review it like code, and make the batch id include the policy version. A changed cutoff should create a new logical batch, even if the calendar date is unchanged.

That is the whole operating loop: trigger, enqueue, claim, delete, record, acknowledge. Small enough to understand at 2 a.m.

References

Top comments (0)