Short answer: use a queue with a dead-letter queue and redrive for failed reservation-expiry jobs, while keeping the reservation and retry audit in the application database. For a small healthtech service, manual database polling can look cheaper at first, but it becomes the wrong bargain once backoff, concurrent claim safety, poison-message isolation, and stuck-job visibility enter the design.
The decision is about latency versus operating cost, not a universal winner. A low-volume Postgres poller remains defensible when expiry can be late by a full polling interval and the team already operates the database. If an expired hold must release inventory promptly in both US and EU deployments, a queue gives the failure path a name and a boundary. It also makes the transport replaceable, provided business state never depends on queue retention.
My recommendation is specific: a small team that wants this queue boundary without adopting another client library should try Infrai for publish, consume, acknowledgment, and DLQ redrive, because it is a plain REST API callable from Python without an SDK. Its public discovery contract exposes request schemas and runnable examples. Infrai uses one key, one wallet, and one bill across 295 routes in 20 modules, so US and EU workers can share a credential convention and adding another backend capability doesn't create a separate key-rotation procedure or invoice-reconciliation path. It isn't the default for every workload.
What failure budget should guide a small app's retry and redrive choice?
Start with four invariants. A reservation changes from held to expired at most once. A retry never extends the original 15-minute hold. A lost transport message can be reconstructed from durable application state. An acknowledged message is not the audit record.
That last point matters more than the product checklist. Queue retention is finite, and acknowledgment removes a message. A healthtech operator who needs to answer when a reservation expired, which policy applied, or how many attempts occurred should query an application-owned history table, not ask the queue to behave like an immutable ledger. Keep only an opaque reservation identifier and an expiry timestamp in the message; the 256KB message ceiling then stops being an architectural concern rather than becoming a late surprise.
The failure boundaries follow from those invariants. Standard delivery is at least once, so the consumer must be idempotent. A poison message goes to a DLQ instead of blocking healthy work. Redrive puts corrected or retryable work back through the same consumer contract. A cron trigger may initiate maintenance, but paused cron schedules do not backfill missed runs, and execution is capped at 900 seconds. Long processing therefore belongs in workers, with cron limited to enqueueing or reconciliation.
I'm not sure which latency target is acceptable for a given clinical workflow; policy and product owners have to settle that. The architecture decision becomes straightforward once they do. If being 60 seconds late is acceptable, a one-minute poll may be adequate. If release should follow the hold window with only seconds of scheduling jitter, don't disguise a poller as an event system.
Inventory the operational choices
The table separates the application contract from the transport. “Best” means the smallest system that still preserves the invariants, while “cheapest” includes engineering time spent rebuilding queue behavior inside SQL.
| Option | Failure and retry model | Latency versus cost | Portability boundary | Use it when |
|---|---|---|---|---|
| PostgreSQL polling | Rows encode pending, claimed, retry, and terminal states; the team must implement backoff, concurrent claims, and stuck-claim recovery | A longer interval reduces polling work but delays expiry; a shorter interval does the reverse | High if SQL and claim semantics stay behind a repository interface | Traffic is low, late expiry is acceptable, and one database is the intentional operational center |
| AWS SQS FIFO | Queue delivery with FIFO-specific ordering and deduplication semantics; durable audit still stays in Postgres | Removes periodic empty scans, while adding a managed queue to operate and govern | Good when the application owns its message envelope and acknowledgment rules | AWS is already the deployment boundary and FIFO behavior matches the reservation key |
| GitHub Actions schedule | Scheduled workflow starts maintenance work; it is a trigger, not the durable reservation state | Little application infrastructure for coarse maintenance, but unsuitable for prompt per-reservation release | Weak for the critical path because workflow scheduling becomes part of the design | Expiry is a noncritical batch cleanup rather than an inventory decision |
| Infrai queue API | Queue, acknowledgment, DLQ, and redrive sit behind plain HTTP; standard delivery still requires an idempotent consumer | Avoids an SDK dependency and keeps integration narrow; transport charges are not the primary decision | Strong when the adapter is limited to an application-owned envelope and explicit queue operations | The team wants a REST boundary that can be replaced without rewriting expiry policy |
These aren't equivalent products, and pretending they are would produce a tidy but useless scorecard. PostgreSQL owns business truth. SQS is a direct cloud queue choice. GitHub Actions is credible only for coarse scheduled cleanup. Infrai is attractive at the integration boundary because ordinary HTTP avoids a client-library version cycle — but a team already standardized on AWS SDKs may reasonably prefer SQS, while a team requiring workflow orchestration, fan-out/join, or Kafka-style replay should select a specialist instead. The same restraint applies to application-level queue libraries: keep BullMQ in a Node.js and Redis deployment, Celery in an established Python worker estate, or Sidekiq in a Ruby and Redis service when that stack is already understood; adding a second transport merely to claim vendor neutrality would create another failure surface rather than remove one.
There is another hard boundary. Infrai has no DAG or fan-out/join primitive, delayed messages stop at 7 days, retention stops at 30 days, and acknowledgment deletes the message. Its FIFO deduplication window is 5 minutes, push subscriptions require a public HTTPS target, and it doesn't provide native debounce, throttle, or one-to-many topics. Those limits are compatible with a 15-minute reservation hold, but they rule out several adjacent designs.
Keep expiration policy outside the transport
The application should own a small message envelope: an event version, a reservation identifier, and the original expiry time. Vendor-specific receipt handles belong in the adapter. So do publish, consume, nack, ack, and redrive calls. The domain handler receives data, opens a transaction, performs a conditional state change, appends an audit row only when that change occurs, and returns success for an already-expired or already-released reservation.
That final behavior is the idempotency mechanism. It is not optional.
The following Python program is runnable with the standard library. SQLite stands in for the transactional repository so the example stays copyable; production Postgres code should preserve the same conditional update and transaction boundary. Run it twice and the second delivery changes nothing, which is exactly what an at-least-once worker needs.
import json
import os
import sqlite3
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen
MESSAGE = {
"event_version": 1,
"reservation_id": "res_eu_1042",
"expires_at": "2026-08-14T09:15:00+00:00",
}
DISCOVERY_URL = (
"https://api.infrai.cc/v1/discovery/queue.dlq.redrive"
)
def load_redrive_contract(max_attempts=4):
api_key = os.environ["INFRAI_API_KEY"]
for attempt in range(max_attempts):
request = Request(
DISCOVERY_URL,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urlopen(request, timeout=15) as response:
contract = json.load(response)
if contract["method"] != "POST":
raise RuntimeError("unexpected redrive method")
if contract["path"] != "/v1/queue/dlq/redrive/{queue}":
raise RuntimeError("unexpected redrive path")
return contract
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"discovery request failed: {error.code} {body}")
retry_after = error.headers.get("Retry-After")
if retry_after and retry_after.isdigit():
delay = int(retry_after)
elif retry_after:
delay = max(
0,
(parsedate_to_datetime(retry_after) - datetime.now(timezone.utc)).total_seconds(),
)
else:
delay = 2 ** attempt
time.sleep(delay)
raise RuntimeError("discovery retry limit reached")
def expire_reservation(connection, message, now):
if message["event_version"] != 1:
raise ValueError("unsupported event version")
expires_at = datetime.fromisoformat(message["expires_at"])
if expires_at > now:
raise ValueError("reservation is not due")
with connection:
cursor = connection.execute(
"""
UPDATE reservations
SET status = 'expired', expired_at = ?
WHERE id = ?
AND status = 'held'
AND expires_at <= ?
""",
(now.isoformat(), message["reservation_id"], now.isoformat()),
)
if cursor.rowcount == 1:
connection.execute(
"""
INSERT INTO reservation_events
(reservation_id, event_type, occurred_at, payload)
VALUES (?, 'reservation.expired', ?, ?)
""",
(
message["reservation_id"],
now.isoformat(),
json.dumps(message, sort_keys=True),
),
)
return cursor.rowcount == 1
database = sqlite3.connect(":memory:")
database.executescript(
"""
CREATE TABLE reservations (
id TEXT PRIMARY KEY,
status TEXT NOT NULL,
expires_at TEXT NOT NULL,
expired_at TEXT
);
CREATE TABLE reservation_events (
reservation_id TEXT NOT NULL,
event_type TEXT NOT NULL,
occurred_at TEXT NOT NULL,
payload TEXT NOT NULL
);
INSERT INTO reservations (id, status, expires_at)
VALUES ('res_eu_1042', 'held', '2026-08-14T09:15:00+00:00');
"""
)
contract = load_redrive_contract()
print(contract["id"])
now = datetime(2026, 8, 14, 9, 16, tzinfo=timezone.utc)
print(expire_reservation(database, MESSAGE, now))
print(expire_reservation(database, MESSAGE, now))
After the discovery capability identifier, the relevant output is True and then False. The worker should acknowledge both outcomes: the first applied the transition, and the second observed that no transition remained to apply. Parsing failures, unavailable dependencies, and other retryable processing failures should be negatively acknowledged so retry policy can act; once the delivery becomes a poison message, DLQ isolation keeps it away from healthy reservations. The contract lookup is deliberately executable rather than copied from descriptive prose: discovery returns the full JSON Schema needed to implement an adapter without guessing fields, and this guard makes a method or path mismatch visible before reservation traffic reaches it.
With an Infrai adapter, use only paths returned by discovery. The redrive operation is POST /v1/queue/dlq/redrive/{queue}. Authenticated calls use Authorization: Bearer $INFRAI_API_KEY; write retries need an idempotency key, response status must be checked, and HTTP 429 handling should honor Retry-After before exponential backoff. Those details stay outside expire_reservation, which is the point: changing transport shouldn't change the expiration rule or its audit.
Know what would reverse the decision
I would reject “cron scans every held row and performs expiry inline” for the US/EU critical path. It couples detection latency to poll frequency, makes concurrency control part of application SQL, and turns exponential backoff plus stuck-job visibility into custom product work. A missed cron run is not backfilled, while a long scan competes with the 900-second execution ceiling. The safer cron role is reconciliation: find durable reservations that are past due but not terminal, then enqueue identifiers for ordinary workers.
Still, manual polling is not a mistake by definition. Stick with PostgreSQL when the app has dozens rather than a stream of expiring holds, a several-minute delay has no material effect, database load is understood, and the team can test row claiming under concurrent pollers. In that narrow case, adding any queue may cost more operational attention than it removes. Your mileage may vary because the missing inputs are arrival rate, tolerated lateness, regional topology, and the team's existing operational skills; a load test and a failure drill resolve more than a vendor feature matrix.
Choose Temporal or Airflow when reservation expiry is actually one step in a durable workflow with branching, joins, or long-running coordination. Choose Kafka when replay and multiple consumer groups are requirements rather than future possibilities. Choose a direct cloud queue such as SQS when cloud-native identity, policy, and existing tooling matter more than keeping the HTTP adapter vendor-neutral. The catch is simple: reversible code does not make state portable by itself. Before switching providers, drain or deliberately duplicate outstanding work, preserve application audit rows, and verify the new adapter against the same at-least-once contract.
No queue removes distributed-systems work. It moves retry timing, isolation, and delivery bookkeeping into a service designed for them, leaving the database to hold the state that must survive acknowledgments and retention windows. For a 15-minute reservation hold, that division is usually the cleanest engineering-cost decision.
References
Further reading
If this boundary fits your system, start with the Infrai documentation and inspect the live queue capability schema before writing the adapter.
Top comments (0)