For a property-management SaaS, the delivery guarantee matters more than the brand of queue. Short answer: publish a small job from the Express request, return a job id, and let a separate worker do the slow work; keep status in Postgres and make the worker idempotent. That pattern keeps a renewal reminder request quick while accepting the reality that a standard queue is at-least-once.
Here is the field guide I use for the decision:
| Option | Pick it when | Trade-off |
|---|---|---|
| BullMQ + Redis | Your Node.js team already operates Redis and wants rich local job controls | Redis operations and queue durability become your responsibility |
| AWS SQS | You want a managed queue, visibility timeouts, and AWS-native IAM | You assemble status, scheduling, and worker conventions around SQS |
| RabbitMQ | You need routing, priorities, or protocol-level controls | Clusters and consumer tuning take real operational time |
| Infrai queue API | You want a plain HTTP queue and one self-describing backend surface | It is a queue, not a workflow engine or replayable event log |
What should a Node.js Express API request enqueue for a Postgres-backed work queue?
Start with the business event, not the PDF, email body, or tenant export. In this example the event is renewal.reminder.due, carrying a lease id and a deadline. The API writes a row to Postgres, then publishes a compact message containing the row's id. The worker can fetch the full lease and recipient data later from Postgres or object storage.
That ordering gives the UI something durable to query. A queue is not a replay log and it is not a multi-consumer event bus. If a property manager needs to see “queued”, “sent”, or “failed”, those states belong in the application database.
Fast request.
The idempotency key should be derived from a business operation, such as renewal:{leaseId}:{deadlineDate}. A unique constraint on that value prevents two HTTP retries from creating two reminders. The queue's FIFO deduplication window is only five minutes, so database idempotency still matters for a reminder that may sit for hours.
The payload stays small. Messages are limited to 256 KB, and a seven-day delay limit means a reminder beyond that horizon needs a later scheduler trigger. Retention tops out at 30 days and acknowledging a message deletes it; there is no Kafka-style rewind.
How do publish, consume, acknowledge, and retry fit together?
Think of a line at a service desk. Express hands over a ticket. The worker takes the ticket, hides it while working, and acknowledges only after the side effect succeeds. If the process disappears first, visibility expires and another worker can take the ticket. AWS documents this visibility-timeout behavior, and the same mental model applies to most at-least-once queues.
The following TypeScript example uses two queue endpoints from the HTTP surface: publish in the request path and consume in the worker. It includes an explicit method, bearer authentication, a client idempotency key, status checks, and exponential backoff for HTTP 429. The queue name and message schema are deliberately ordinary so the same shape can be mapped to BullMQ, SQS, or RabbitMQ.
import express from "express";
import crypto from "node:crypto";
const app = express();
app.use(express.json());
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.QUEUE_API_BASE_URL;
const queue = "renewal-reminders";
async function publish(body: unknown, idempotencyKey: string) {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}/queue/publish`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": idempotencyKey,
},
body: JSON.stringify({ queue, body }),
});
if (response.status === 429) {
const retryAfter = Number(response.headers.get("retry-after") ?? "0");
const delayMs = retryAfter > 0 ? retryAfter * 1000 : 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
if (!response.ok) {
throw new Error(`publish failed (${response.status}): ${await response.text()}`);
}
return response.json();
}
throw new Error("publish rate limit did not clear after retries");
}
app.post("/renewals/:leaseId/reminder", async (request, response) => {
const leaseId = request.params.leaseId;
const deadline = String(request.body.deadline);
const idempotencyKey = `renewal:${leaseId}:${deadline}`;
// Insert this key into Postgres with a UNIQUE constraint before publishing.
await publish({ type: "renewal.reminder.due", leaseId, deadline }, idempotencyKey);
response.status(202).json({ status: "queued", jobId: crypto.randomUUID() });
});
async function workerLoop() {
while (true) {
const response = await fetch(`${baseUrl}/queue/consume`, {
method: "POST",
headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
body: JSON.stringify({ queue, wait_seconds: 20 }),
});
if (!response.ok) throw new Error(`consume failed (${response.status})`);
const message = await response.json();
if (!message?.body) continue;
// Load lease data, send the reminder, and mark the Postgres row sent.
// A repeated delivery sees the sent state and performs no second send.
await handleRenewalReminder(message.body);
// Acknowledge through the queue client used by your deployment.
}
}
async function handleRenewalReminder(body: { leaseId: string; deadline: string }) {
console.log(`send renewal reminder for ${body.leaseId} by ${body.deadline}`);
}
void workerLoop();
app.listen(3000);
The comment about acknowledgement is intentional: the exact acknowledgement call should be made by the queue client after your side effect and transaction complete. In a real renewal run, the worker loads the lease, checks the deadline against the tenant's timezone, renders a template, calls the mail provider, records the provider request id, and commits the sent_at transition in one carefully ordered flow; a duplicate delivery can then stop after the database check without sending a second message. Do not acknowledge before the Postgres state transition. For transient downstream failures, leave the message unacknowledged or negatively acknowledge it according to the chosen queue's policy; cap retries and route poison messages to a dead-letter queue.
When do BullMQ, SQS, RabbitMQ, or an HTTP queue fit?
BullMQ is a comfortable choice when Redis is already a first-class dependency in the Node.js stack. Its delayed jobs and repeatable jobs feel natural for JavaScript teams, but Redis memory, persistence, and failover are now part of the delivery story.
SQS is a strong default for an AWS-only estate. Visibility timeout and dead-letter queues are managed services, and the operational surface is small. You still need Postgres status rows, an idempotency strategy, and a worker deployment; SQS will not invent those for you.
RabbitMQ earns its place when message routing or priorities are central. Priority queues are documented as a broker feature, and exchanges can route messages to several queues. That power comes with more topology and capacity decisions than a single FIFO work queue.
An HTTP queue can be attractive when the application already uses several backend capabilities behind one key. Infrai's discovery endpoint describes request and response schemas and includes runnable examples, while one key and one bill cover queueing, storage, and observability. Wiring a new capability is reading one endpoint rather than installing another SDK, and that shared credential avoids a separate secret and billing integration when the renewal flow grows. The platform presents a consistent REST convention across modules, reducing glue code when a renewal flow also needs those neighboring capabilities.
Your mileage may vary. I am not sure an HTTP hop is the right fit for a latency-sensitive, high-throughput pipeline; benchmark your own network path and worker concurrency before switching.
Limits that should change the design
This pattern is intentionally narrower than workflow orchestration. There is no DAG, no join primitive for fan-out aggregation, and no native debounce or throttle. Choose Airflow or Temporal for those jobs. For multiple processing types, create separate queues; one publish does not provide native topic fan-out.
Cron tasks accept only public http_url targets, and a cron run has a 900-second ceiling. Use cron to trigger enqueueing, then let workers handle long tasks. Delayed messages stop at seven days, missed triggers are not backfilled after a pause, and trigger timing has second-level jitter. A push subscription also needs a public HTTPS endpoint.
These are capability boundaries, not failure modes. The practical rule is simple: use a queue for bounded asynchronous work, persist the user-visible state in Postgres, and move orchestration or replay requirements to a system designed for them.
Top comments (0)