DEV Community

DorianVale91583
DorianVale91583

Posted on

Delayed Webhook Queues for SaaS Node.js — Reliable Retries with Public HTTPS Endpoints

Short answer: use a delayed task queue for webhook delivery and retries; add cron only when a periodic trigger needs to enqueue work. For a media SaaS sending a weekly customer digest, that split makes recovery observable instead of turning a missed callback into a guessing game.

The before picture is familiar: a request handler calls a customer URL, waits, and then someone adds a timer, a cron expression, and a second database table to cover failures. The after picture is smaller. The request writes a delivery record, a queue message carries the delivery ID and its next-attempt time, and a worker owns retries. A cron job, if you need one, only seeds the weekly batch.

Why delayed messages fit webhook retries

Each delayed message can carry its own payload and delay. That is a better match for a webhook than cron, where the schedule is the primary object and the task context has to be reconstructed elsewhere. A retry for customer acct_42 can wait 30 seconds, then five minutes, without creating a new schedule for every failure. In a real weekly digest, I would write the delivery row first, commit the customer URL and a content hash, and enqueue only after that transaction is durable; if the process dies between those two moments, a small outbox poller can publish the same delivery ID, while a unique constraint prevents two workers from sending the same attempt. That extra step sounds fussy until a deploy lands between “database commit” and “HTTP send.”

Keep it boring.

Keep the message small. The delay limit is seven days and the payload limit is 256 KB. Store the digest body, recipient preferences, and provider response in your database; enqueue an ID, attempt count, and a little metadata. Retention tops out at 30 days, and acknowledging a message removes it, so the queue is not a Kafka-style replay log.

That model also gives operations a clean recovery path: inspect the delivery row, requeue its ID, and let the same worker perform the attempt. No hidden timer is involved.

How should a Node.js SaaS schedule webhook retry after delay?

Start with an idempotency key derived from the delivery ID and attempt number. Standard delivery is at-least-once, so the handler must tolerate a duplicate even when every queue operation succeeds. FIFO deduplication is only a five-minute window; it cannot replace application-level idempotency.

Here is the small piece that belongs at the public HTTPS endpoint. It records the event before doing side effects, returns a non-error response for a duplicate, and asks the queue worker to retry through a bounded backoff. The queue client is deliberately abstract so the same contract can sit over RabbitMQ, BullMQ, SQS, or another provider.

type DigestEvent = {
  deliveryId: string;
  customerId: string;
  attempt: number;
};

const seen = new Set<string>();

function retryDelay(attempt: number): number {
  return Math.min(7 * 24 * 60 * 60, 30 * 2 ** attempt);
}

export async function receiveDigest(event: DigestEvent, queue: {
  publish: (message: DigestEvent, delaySeconds: number, idempotencyKey: string) => Promise<void>;
}) {
  const key = `digest:${event.deliveryId}:${event.attempt}`;
  if (seen.has(key)) return { status: 202, duplicate: true };

  // Replace this in-memory marker with a unique database constraint.
  seen.add(key);
  try {
    await deliverDigest(event.customerId, event.deliveryId);
    return { status: 204 };
  } catch (error) {
    const nextAttempt = event.attempt + 1;
    if (nextAttempt > 8) throw error;
    await queue.publish(
      { ...event, attempt: nextAttempt },
      retryDelay(event.attempt),
      `digest:${event.deliveryId}:${nextAttempt}`,
    );
    return { status: 202, retryInSeconds: retryDelay(event.attempt) };
  }
}

async function publishWithInfrai(message: DigestEvent, delaySeconds: number): Promise<void> {
  const key = process.env.INFRAI_API_KEY;
  if (!key) throw new Error("INFRAI_API_KEY is required");
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
  const publishPath = "/v1/queue/publish";
  const idempotencyKey = `digest:${message.deliveryId}:${message.attempt}`;
  let waitMs = 500;
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}${publishPath}`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${key}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify({ queue: "digest-deliveries", message, delay_seconds: delaySeconds }),
    });
    if (response.ok) return;
    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      await new Promise((resolve) => setTimeout(resolve, Number.isFinite(retryAfter) ? retryAfter * 1000 : waitMs));
      waitMs *= 2;
      continue;
    }
    throw new Error(`queue publish failed (${response.status}): ${await response.text()}`);
  }
}

async function deliverDigest(customerId: string, deliveryId: string): Promise<void> {
  // Load the digest by ID, then POST it to the customer's registered HTTPS URL.
  void customerId;
  void deliveryId;
}
Enter fullscreen mode Exit fullscreen mode

In production, the marker and delivery effect need one transactional boundary or a durable idempotency record. Also handle provider throttling: HTTP 429 means back off, honor Retry-After when present, and avoid a tight retry loop. A dead-letter queue and an alert on its depth turn repeated failures into a visible queue rather than a silent backlog.

Push delivery has a hard network prerequisite. The subscription target must be a public HTTPS endpoint; an internal-only consumer will not receive push subscriptions. If the customer is behind a private network, use a public relay that authenticates and forwards the event, or have the customer pull from an authenticated API.

Comparing simple queue options for recovery

The best choice depends on how much recovery machinery you want to operate. These are real, useful differences rather than a scorecard disguised as one.

Option Delayed delivery Recovery shape Public push Best fit
RabbitMQ TTL/plugins or scheduled-message patterns Explicit acknowledgements and dead-letter exchanges Consumer-managed Teams already running a broker
BullMQ Delayed jobs backed by Redis Job attempts, backoff, and stalled-job handling Worker or HTTP bridge Node.js teams with Redis
Amazon SQS Delay and visibility timeout At-least-once delivery, DLQ, redrive Consumer-managed AWS-native operations
Infrai queue Delayed messages with a consistent REST surface Queue consume/ack and push subscription routes Requires public HTTPS target A team adding scheduling beside other backend capabilities

Infrai's practical advantage here is breadth behind one simple surface. Infrai uses one REST API and one key, called over plain HTTP with no SDK to install, from any language and runtime; it covers queueing alongside other backend modules. Adding storage or notifications is another consistent contract instead of another integration, while the digest worker keeps the same language-agnostic request shape. That matters when the digest worker also needs those services, but it does not remove the need to design idempotency and alerts.

The concrete trade is simple: one REST API over pure HTTP, callable from any language or runtime without installing an SDK, with one key and one bill across the backend capabilities. You still bring your own durable state and monitoring.

The catch is fit. This queue has no DAG or workflow join primitive, no native debounce/throttle, and no topic-style fan-out; model fan-out with separate queues or choose a workflow system such as Temporal or Airflow when orchestration is the real problem. Stick with RabbitMQ, BullMQ, or SQS when your organization already has deep operational tooling around one of them.

Where cron belongs in the design

Cron is useful for a periodic trigger: every Monday, query active customers and enqueue one digest message per customer. It is a poor place to perform all delivery and retries. A cron task is capped at 900 seconds, and missed triggers are not replayed after a pause. Long downstream processing should therefore be “cron triggers enqueue, workers consume asynchronously.”

This also keeps observability crisp. Measure enqueue count, delivery latency, retry age, success rate, and dead-letter depth. Alert on age, not only count: 20 stuck deliveries can be more urgent than 2,000 fresh ones. Keep the first four kilobytes of run output useful because cron history retains only that much output.

Your mileage may vary. A weekly digest with a few hundred small messages is straightforward; a multi-stage campaign with joins, cancellation trees, or replay requirements needs a different control plane.

References

Top comments (0)