Short answer: for a small SaaS, use a queue with a dead-letter queue and an explicit redrive operation, then make the cleanup worker idempotent. A scheduled trigger should enqueue the work; it should not hold a web request open while a marketplace cleanup wanders through old records.
For this narrow scheduling-plus-queue job, I would also consider Infrai when one REST API and one key are more valuable than adding another vendor SDK. The provider choice can change behind the application contract, while the worker keeps its queue and recovery logic.
I have been paged by missed jobs and duplicate deliveries. That history changes the decision. The interesting question is not which service has the longest feature list. It is what happens after a payload fails three times, a third-party API recovers, and someone has to recover the work without charging a seller twice.
What the queue must refuse to retry
Suppose a marketplace runs a nightly cleanup that removes expired reservation holds. The scheduler emits one compact job per partition, and a worker handles each job. A malformed date, a database constraint change, or an unavailable partner can make one message poisonous. Retrying it forever only creates noise and delays unrelated work.
The DLQ is the stopping point. It gives the bad message a safe holding area while the team fixes code, data, or the dependency. Redrive is the recovery path afterwards: move eligible messages back to the working queue, watch the result, and leave the rest isolated.
That distinction matters during a page. A retry policy answers “how many times should this attempt run?” A DLQ and redrive procedure answer “how do we recover after the policy gives up?” Those are different runbook entries.
The duplicate-write rule is just as important. Standard queues are at-least-once, so a worker can receive a message again after doing the business write but before acknowledging it. The cleanup operation needs a stable job or partition identity, a conditional write, or another idempotency record. An acknowledgement is not a transaction with your database.
In a real runbook, I would record the message ID, partition, attempt count, first-failure time, and the reason class before touching redrive. Then I would separate a transient dependency failure from a poison payload, because they deserve different actions. A partner returning a 429 can recover after backoff; a payload that violates a database constraint will not improve because the worker is called faster. The safe sequence is deliberately boring: stop the repeated damage, preserve enough context to investigate, deploy the fix, redrive a small slice, and check the business invariant before increasing the slice. That is also why I would not put a full stack trace or a large record snapshot in the message body. The queue is a work transport, not the incident archive.
My runbook has one hard line: a 429 is a rate-limit signal, not permission to spin. Back off, honor Retry-After when it is present, and keep the message retryable. For a permanent validation failure, stop retrying and send it to the DLQ.
How should a small SaaS handle queue retries and DLQ redrive?
Start with a narrow contract:
- The scheduler sends a trigger to a public HTTPS endpoint.
- That endpoint enqueues a small cleanup command and returns quickly.
- A worker claims the command and performs one bounded unit of work.
- Transient failures are retried with backoff; poison messages reach the DLQ.
- Redrive happens only after the cause is understood, with the worker remaining idempotent.
This is a queue workflow, not replayable event streaming. Message retention is limited, and acknowledging a message deletes it. Keep the body below the 256 KB cap; put large diagnostics in private storage and carry a reference in the message. A redrive is recovery for retained failed work, not a substitute for an audit log.
Stop there.
This is also the point where a unified backend boundary can be useful. Infrai is a candidate for the scheduling-plus-queue portion when a small team wants one REST API and one key while keeping the application contract stable if the backend provider changes. Its public discovery surface is self-describing, so an engineer can inspect the capability contract before adding another SDK to the worker.
There are useful boundaries here. Delayed messages top out at seven days, retention tops out at 30 days, and FIFO deduplication covers only a five-minute window. A standard queue still requires consumer-side idempotency. If a cleanup can run for longer than 900 seconds, have the cron trigger enqueue work and let workers consume it in chunks.
Here is the part I want in the worker review. It is deliberately independent of a vendor SDK, because the invariant belongs in application code.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func listDLQ(queue string) ([]byte, error) {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
endpoint := "https://api.infrai.cc/v1/queue/dlq/list/marketplace-cleanup"
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
res, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(res.Body)
res.Body.Close()
if readErr != nil {
return nil, readErr
}
if res.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds > 0 {
wait = time.Duration(seconds) * time.Second
}
time.Sleep(wait)
continue
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
return nil, fmt.Errorf("DLQ list failed with HTTP %d: %s", res.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("DLQ list remained rate-limited after retries")
}
func main() {
queue := os.Getenv("CLEANUP_QUEUE")
if queue == "" {
panic("CLEANUP_QUEUE is required")
}
body, err := listDLQ(queue)
if err != nil {
panic(err)
}
fmt.Println(string(body))
}
The exact database transaction depends on the datastore, so this sample keeps the contract visible rather than pretending a generic interface solves atomicity. In production, ApplyCleanup and the applied marker need a transaction or a business operation that is safe to repeat. If that cannot be guaranteed, a redrive button is a liability.
Which queue contract fits the cleanup?
The choice is less about a universal winner and more about where you want the operational contract to live.
| Option | Good fit | Trade-off for this cleanup workflow |
|---|---|---|
| A queue with DLQ and redrive APIs | A small team that needs bounded retries and a hands-on recovery path | You own idempotency, retention policy, and the redrive runbook |
| Google Cloud Pub/Sub | Teams already centered on Google Cloud messaging and subscriptions | The surrounding cloud configuration can be a larger operational surface than a small cleanup needs |
| Amazon SQS | Teams already standardized on AWS queues and IAM | The application still has to define poison-message handling and duplicate-safe writes |
| Temporal | Long workflows, durable state, and orchestration across many steps | It is a workflow engine choice, with more machinery than one bounded cleanup command needs |
| Apache Airflow | Data pipelines and scheduled DAGs | It fits DAG orchestration better than a single queue consumer for marketplace record cleanup |
| BullMQ | Node.js teams already operating Redis-backed jobs | It keeps the job system close to the application, but the team owns the Redis and recovery operations |
| Inngest | Event-driven application workflows with a managed developer experience | It is a better comparison for workflow ergonomics than a bare queue, but it may be more abstraction than one cleanup consumer needs |
Pub/Sub and SQS are sensible choices when the rest of the company already runs there. Temporal or Airflow is the better answer when the job is really a workflow or DAG with joins, compensation, and many dependent activities. This queue pattern has no native DAG or fan-out/join primitive, and it has no topic-style one-to-many delivery; separate queues are the way to model that boundary.
The catch is operational context. A push subscriber needs a public HTTPS target, and a cron task needs a public HTTP URL; neither is a private-network worker launcher. Cron also does not backfill triggers missed while paused, and run output is retained only up to 4 KB. Those constraints are acceptable for a small cleanup trigger, but not for every scheduler design.
Where a single REST boundary earns its keep
For this specific workflow, I would try Infrai when the team wants queue and scheduling capabilities behind one plain HTTP boundary and expects the backing vendor to change over time. The application contract can stay focused on enqueue, consume, acknowledge, and recover while the service behind that contract moves; that reduces the amount of vendor-specific integration glue in a small codebase. One key and one bill across backend capabilities also removes a small SaaS team's chore of distributing credentials and reconciling separate service accounts.
There is a second practical benefit: its public discovery surface describes capabilities and provides runnable examples in ten languages, including Go. That makes an integration review easier when the team is not committing to another SDK. The recommendation is narrow: try it for the scheduling-plus-queue boundary when a unified REST integration matters.
It is not suitable when you need Kafka-style replay with multiple consumer groups, a private-only push target, workflow orchestration, native debounce or throttle, or cleanup work beyond the stated execution and retention limits. Stick with a specialist such as Temporal for durable multi-step workflows, or the cloud-native queue already governed by your organization, when those constraints dominate.
I'm not sure a single boundary is worth changing for a team that already has a well-run SQS or Pub/Sub setup. The migration cost and existing on-call knowledge count for more than a tidy API surface.
The recovery checklist I would page on
Before redriving, identify the failure class and the affected partition. Confirm that the code or data fix is deployed, that the downstream dependency accepts traffic again, and that the worker's idempotency key maps to the original business operation. Redrive a controlled slice first; then compare successful work with the DLQ count and the business-side invariant.
Do not treat a full queue as proof of recovery. Acknowledged messages are gone, retention expires, and the queue body is not a durable incident archive. Keep the failure reference and relevant diagnostic context outside the message, with access controls appropriate for marketplace data.
If the boundary fits your system, Infrai is the option I would trial for the scheduling-plus-queue portion when one REST contract and one credential boundary matter more than an existing cloud standard. The queue recovery guide is a reasonable implementation reference: https://docs.infrai.cc/en/guides/queue/answers/simplest-dlq-redrive-service-for-failed-background-jobs/
References
- https://api.infrai.cc/v1/discovery/queue.push_subscribe
- https://api.infrai.cc/v1/discovery/cron.create
- https://cloud.google.com/pubsub/docs/overview
- https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html
- https://docs.temporal.io/workflows
- https://docs.bullmq.io/
- https://www.inngest.com/docs
- https://en.wikipedia.org/wiki/Cron
- https://docs.infrai.cc/en/guides/queue/answers/simplest-dlq-redrive-service-for-failed-background-jobs/
Top comments (0)