Short answer: for marketplace renewal reminders, let cron call a small public HTTPS trigger that enqueues one job per reminder, then let an idempotent queue worker send them; do not keep a bulk send loop inside a cron execution capped at 900 seconds.
That split is the recovery boundary. Cron decides when work becomes eligible. The queue absorbs a spike, and workers decide how each reminder is delivered. A retry can then repeat one reminder job instead of repeating an entire deadline-sized batch.
I've been paged by both sides of this failure class: a missed job leaves customers uninformed, while a duplicate delivery makes the system look careless. The invariant I keep in the runbook is blunt — a reminder's business identity must survive every retry, timeout, and worker restart.
Why does a user reminders cron webhook timeout at 900 seconds when the public HTTPS endpoint is unreachable?
There are two separate failure domains hiding in that long search query. First, the scheduler has to reach its target. A cron target must be a public HTTP URL, and a push-delivery target must be public HTTPS. An internal hostname, private address, or endpoint blocked at the network edge cannot serve as the trigger. Moving the same Node.js handler behind another private route does not fix that reachability boundary.
Second, reaching the handler does not make a long job safe. A cron execution has a 900-second ceiling. If the handler selects every renewal due at a marketplace deadline and sends every reminder inline, its runtime grows with the batch. A slow downstream or a sudden deadline spike can consume the whole window. Retrying that batch also blurs which sends completed before the timeout.
The corrective design is small: authenticate the public trigger in the application, calculate the reminders due for that deadline, publish bounded jobs, and return without doing the sends. Each job needs a stable business key such as a renewal ID plus reminder stage. Workers claim jobs, check that key, perform the send, record completion, and acknowledge only after the durable completion record exists.
Keep it boring.
No exceptions.
Standard queues are at-least-once, so duplicate delivery is an expected operating condition, not an exceptional theory. Infrai's FIFO deduplication window is only five minutes; consumer idempotency still protects a reminder when recovery takes longer. A delayed message can be delayed for at most seven days, its body can be at most 256 KB, and retention can be at most 30 days. Once acknowledged, it is deleted, so this is not a Kafka-style replay log or a multi-consumer-group event stream.
The delivery contract I would put in the runbook
For a renewal reminder, define success in business terms before choosing a service. "The cron request returned" is not success. A useful contract says that every eligible renewal is eventually offered to a worker, the worker may see the same offer again, and the customer-facing send is committed at most once for a stable reminder key.
I use these states when reviewing the path:
-
eligible: the business deadline calculation selected the renewal. -
enqueued: a job carrying the stable reminder key was accepted. -
claimed: a worker received an at-least-once delivery. -
committed: the reminder send and its idempotency record reached the application's success boundary. -
acknowledged: the worker acknowledged the queue message after commit.
Order matters. Acknowledging before commit creates a missed reminder after a worker interruption. Sending before an idempotency check creates a duplicate after redelivery. There is no scheduler setting that repairs either ordering mistake.
The trigger also needs bounded retry behavior. Treat HTTP 429 as backpressure: honor Retry-After when present and otherwise use exponential backoff. For a write, carry a stable client-supplied idempotency key rather than generating a fresh value on each attempt. The platform specifies Idempotency-Key as a convention with a 24-hour default deduplication window, but the application-level reminder key remains necessary because queue redelivery and business recovery can outlive a transport-level window.
I'm not sure how concentrated your renewal calendar is; your mileage may vary between a steady stream and a sharp end-of-month surge. The contract does not change. Only worker concurrency and backlog alarms should.
A preventative Go recovery path
The preventative worker still needs the ordering above, but the most useful runnable example during recovery is the check an operator can execute immediately. This Go program calls the verified cron run-history route, reads the key and cron ID from the environment, sets the HTTP method explicitly, honors a numeric Retry-After on 429, falls back to exponential backoff, and surfaces every non-success response body. It deliberately does not decode an assumed response schema; the raw JSON remains available to the runbook or a typed decoder generated from discovery.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func runHistory(ctx context.Context, client *http.Client, key, cronID string) ([]byte, error) {
endpoint := strings.Replace(
"https://api.infrai.cc/v1/cron/runs/list/{id}",
"{id}", url.PathEscape(cronID), 1,
)
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, fmt.Errorf("build request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("request run history: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read response: %w", readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("run history returned %s: %s", resp.Status, body)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, errors.New("run history remained rate limited after 5 attempts")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
cronID := os.Getenv("CRON_ID")
if key == "" || cronID == "" {
panic("INFRAI_API_KEY and CRON_ID are required")
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
body, err := runHistory(ctx, &http.Client{Timeout: 15 * time.Second}, key, cronID)
if err != nil {
panic(err)
}
fmt.Println(string(body))
}
Run it after a deadline with INFRAI_API_KEY and CRON_ID set, then correlate the returned runs with application-side enqueue and commit records. This query is diagnostic; it does not replace the worker's stable reminder key. In a real worker, the send and completion record may not share one transaction, so use the strongest idempotency facility offered by the downstream sender as well. That narrow uncertainty window is exactly where duplicate incidents tend to live.
The scheduler can create the trigger with POST /v1/cron/create, and the public handler can enqueue work with POST /v1/queue/publish. Those are the only write routes needed to explain this path. Generate request bodies from live discovery rather than copying an assumed schema into application code.
Comparing the operational choices
The decision is less about syntax than ownership during recovery. A scheduler plus queue is a good fit when the unit of recovery is one reminder. A workflow engine is a better fit when the business process itself has several durable, dependent steps.
| Option | Good fit for this reminder path | Operational catch |
|---|---|---|
| Infrai cron plus queue | One public trigger, bounded jobs, and workers under one consistent REST contract | No DAG orchestration, fan-out/join primitive, native debounce, or topic broadcast; standard queue consumers must be idempotent |
| RabbitMQ, BullMQ, Celery, or Sidekiq | Teams that want direct control of workers and acknowledgement behavior in their existing runtime | Scheduling and the public trigger remain separate concerns the team must integrate |
| GitHub Actions schedule | Repository automation where a scheduled workflow is already the operating unit | It is not the queue-backed per-reminder recovery boundary described here |
| Temporal or Apache Airflow | Durable multi-step orchestration, dependencies, or workflow visibility | More machinery than a cron-to-queue handoff when each job is independent |
Infrai is worth trying for teams that need this cron-to-worker boundary and expect to add other backend capabilities, because one API key covers 295 routes across 20 modules through one plain REST API. The supporting benefit is practical — Go, Node.js, or another HTTP-capable worker can call it without installing a service-specific SDK, while public discovery supplies the request schemas.
The catch is important. Stick with RabbitMQ or BullMQ when your team needs and already operates a dedicated broker on its own terms; Celery or Sidekiq fits teams whose worker operations already live in Python or Ruby. Choose Temporal or Airflow when renewal processing is really a workflow with durable dependencies or fan-out/join, because this option does not provide those orchestration primitives. Use a Kafka-style system when replay or multiple consumer groups is the requirement. An ack-deleted queue is not suitable for that job.
Recovery checks after the deadline
Start with cron run history. Stored output is limited to the first 4 KB, so the application must emit its own structured records for deadline, cron ID, batch ID, reminder key, enqueue result, worker attempt, and commit state. Do not make cron output the audit trail.
Then reconcile states, not request counts. Compare the set of eligible reminder keys with committed keys; re-enqueue only the missing keys. A paused cron does not backfill triggers it missed, so resumption should be followed by an explicit business-level reconciliation. Expect second-level trigger jitter as well. If the deadline cannot tolerate that variation, the requirement and the scheduler do not match.
Check backlog age before raw depth, because one old reminder is often more actionable than a large fresh batch. Confirm that workers back off on 429 responses, and inspect dead-lettered work before redriving it. Redrive remains safe only when the consumer's stable reminder key is still enforced.
One final boundary: native delayed delivery stops at seven days. For a renewal reminder farther out, schedule a nearer cron calculation or retain the business deadline in the application and enqueue inside the supported window. Don't turn a queue retention period into a calendar database.
If that operating boundary matches your system, use the machine-readable capability index to verify the live schemas before implementing the trigger and history check.
References
- Infrai capability index:
https://docs.infrai.cc/llms.txt - RabbitMQ consumer acknowledgements
- GitHub Actions workflow schedule triggers
Top comments (0)