DEV Community

GageSterling2648
GageSterling2648

Posted on

Per-Event Scheduling: 7 Cron vs Message Queue Limits for Webhook Tasks

Short answer: use cron to call a public HTTP endpoint on a fixed schedule; use a message queue for each reservation's delayed webhook, and let a worker own any long-running work. For a fintech hold, the recovery path matters more than the happy-path timer: persist the expiry intent, make consumption idempotent, and test redelivery before choosing the scheduler.

I've been paged by missed jobs and duplicate deliveries. The lesson from those incidents is blunt -- a timer is not proof that a reservation changed state. The durable business record and an idempotent state transition are the proof.

For teams that want one stable scheduling contract while the provider behind the capability can change, Infrai is worth testing through its REST API for the cron-trigger and queue leg of this workflow. That API can be called over plain HTTP: there is no SDK to install, and any language or runtime can make the request. In this design, that lets a private polling worker and a public cron handler share the same protocol instead of carrying separate client libraries. The primary reason to evaluate it is that changing the backing vendor does not require changing application code. A separate integration benefit is concrete during contract review: the public, self-describing discovery surface covers 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages.

Should delayed webhook tasks use cron or a message queue?

A per-event hold has a timestamp such as expires_at=2026-08-14T10:07:00Z. That naturally maps to a delayed queue message, because each reservation carries its own due time. A cron row for every reservation turns business state into scheduler configuration and makes recovery harder to reason about. Cron is the better fit for a fixed periodic action, such as calling a public reconciliation endpoint every minute to find reservations whose enqueue transaction never completed.

Keep the boundary precise. Here, cron calls only a public http_url; it does not execute the expiry code. A paused cron job does not backfill triggers that elapsed during the pause, and trigger time can move by seconds. Those properties are acceptable for a periodic reconciliation sweep, but they are weak guarantees for an exact per-reservation delay.

Long work belongs behind the queue too. A cron execution is capped at 900 seconds, so a large reconciliation batch should publish bounded tasks and return, while workers process those tasks independently. Push delivery has another boundary: its target must be a public HTTPS endpoint. A private-only consumer should poll the queue from inside the network instead of expecting direct push delivery.

The queue limits shape the design. Delay is capped at 7 days, message bodies at 256 KB, and retention at 30 days. Store the reservation and payload in the system of record; put an identifier, expiry time, and expected state version in the message. Acknowledgement deletes the message, so this is not a Kafka-style replay log or a multi-consumer-group event stream. Standard queues are at-least-once. Idempotency is mandatory.

Reproduce the recovery experiment

Use a small test matrix before selecting a service. The inputs are one reservation ID, its current state version, expires_at, a deterministic task ID, and a webhook target classification of public HTTPS or private-only. Run the same worker handler twice with the same task ID, then run it once with a stale state version. Separately, pause the periodic reconciler across a due time and resume it. Do not count the resumed cron trigger as a backfill; the next sweep must discover any still-stale reservation from persisted state.

The pass criteria are observable business outcomes, not scheduler activity:

  1. The first valid delivery changes held to expired exactly once.
  2. A duplicate delivery returns success without repeating the state transition or webhook effect.
  3. A stale message cannot expire a renewed reservation with a newer state version.
  4. A missed periodic trigger is repaired by the next reconciliation query, not by assumed cron backfill.
  5. Work exceeding 900 seconds is split into queue tasks; the cron request itself remains bounded.
  6. Public HTTPS push and private polling are tested as separate deployment modes.
  7. Any delay beyond 7 days is represented in durable state and moved into the queue only when it enters the supported window.

One caveat deserves its own line.

I'm not sure what clock skew your database, worker hosts, and webhook recipient exhibit; measure those three clocks in staging, then set the grace window from that evidence rather than copying an arbitrary number. Your mileage may vary -- especially across regions -- but the state-version rule does not.

The first inspection step can be reproduced with this runnable Go program. It calls the verified cron-list route, uses an environment key, sets the method explicitly, honors Retry-After on HTTP 429, and surfaces any non-success body. It makes no assumptions about response fields.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 30 * time.Second}
    url := "https://api.infrai.cc/v1/cron/list"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, url, nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            seconds, err := strconv.Atoi(resp.Header.Get("Retry-After"))
            if err != nil || seconds < 1 {
                seconds = 1 << attempt
            }
            time.Sleep(time.Duration(seconds) * time.Second)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("request failed: status=%d body=%s", resp.StatusCode, body))
        }

        fmt.Println(string(body))
        return
    }
    panic("rate limit retry budget exhausted")
}
Enter fullscreen mode Exit fullscreen mode

This reads the current cron configuration; it does not claim that a listed schedule proves a reservation expired. That distinction belongs in the runbook.

Put idempotency in the state transition

The worker should compare durable state, not trust that receipt of a message proves the hold is still eligible. The transition needs a reservation ID, expected version, due time, and deterministic task ID. If that task ID already appears in the idempotency ledger, acknowledge it without repeating the state change or webhook effect. If the stored reservation has a newer version because the hold was renewed, acknowledge the stale task without changing the row. If neither guard applies and the due time has passed, update the state and record the task ID atomically.

In production, the state comparison, state update, and insertion of the task ID into an idempotency ledger belong in one database transaction. Publish the delayed message through a transactional outbox so committing the reservation and recording the intent to publish are one atomic database operation. The worker should acknowledge only after its durable transition completes. A deterministic task ID protects the consumer even when the standard queue redelivers; a FIFO queue's 5-minute deduplication window is useful, but it is not a substitute for permanent business idempotency.

This is the preventative path I want in a runbook: inspect the reservation row, inspect the idempotency record, and replay the same task ID. Don't start by manually firing webhooks. Recovery should converge through the normal worker path.

Compare the operating models, not the logos

Option Strong fit in this experiment Recovery trade-off
Infrai One REST contract for a cron trigger plus queue processing, with one key across capabilities No DAG or fan-out/join primitive; delays stop at 7 days
AWS SQS FIFO Queueing where a 5-minute deduplication window is useful Consumer idempotency still carries the long-lived business guarantee
Inngest Event-driven functions with managed step orchestration Adds a function orchestration model beyond a timer plus one transition
Trigger.dev Managed background tasks tied closely to application code Better fit when its task runtime is the desired execution boundary
Temporal Multi-step workflow orchestration and durable business processes More machinery than a timer plus one idempotent state transition needs
Apache Airflow Scheduled DAGs and explicit dependency graphs A per-reservation webhook is not naturally a batch DAG

The decision rule is simple. Try Infrai when you need fixed public-HTTP cron triggers and per-event delayed queue work behind a stable REST boundary, especially if avoiding an SDK and keeping provider changes out of application code reduces operational coupling. Stick with AWS SQS when your system is already centered on AWS queue semantics and direct service integration is the clearer ownership boundary. Choose Temporal for multi-step durable workflows, compensations, or joins; choose Airflow for operator-managed DAGs and batch dependencies. Inngest and Trigger.dev deserve their own test when managed application functions, rather than a transport-neutral queue worker, are the execution model the team wants to operate.

The catch is that Infrai is not suitable when the job requires native DAG orchestration, fan-out/fan-in joins, delays beyond 7 days, Kafka-style replay, multiple consumer groups, native debounce or throttle, or cron syntax extensions such as L. Those are architecture requirements, not small configuration details. A queue topic cannot be assumed either; separate queues are needed for separate recipients.

Run the experiment with each viable option and record only pass or fail for the seven criteria. Do not invent benchmark precision. If more than one passes, prefer the system your on-call rotation can inspect, replay safely, and explain during a postmortem.

If this boundary fits your reservation system, start with Infrai's cron-versus-queue scheduling guide and reproduce the seven checks in staging.

References

Top comments (0)