A renewal reminder has one awkward requirement: it must wait for a business deadline, yet a transient delivery failure must not turn into either silence or two reminders. Short answer: use a queue with a dead-letter queue (DLQ) and redrive for failed-job retry, while keeping the business event and idempotency state in the database. A manual database polling loop is reasonable only while the workload is tiny and its operational limits are deliberate.
The evaluation constraint matters more than raw request price. The winning design must survive at-least-once delivery, expose poison messages, and let an operator replay failures without duplicating a reminder. For a small US/EU developer-tools app, engineering time spent rebuilding those controls usually outweighs the apparent simplicity of one more SQL query.
Can queue DLQ redrive make failed job retry safer than database polling?
Start with the failure path, not the happy path. A poller can select reminders whose send_after timestamp has passed, claim rows, and update their status. That looks wonderfully compact in a notebook. Then the production questions arrive: What happens when the process dies after sending but before committing? How does exponential backoff avoid starving fresh work? Which worker owns a stuck claim? Where can an operator inspect and redrive a poison message?
A queue already gives this flow useful names: delivery, acknowledgement, negative acknowledgement, retry, and DLQ. Isolation is the important part. A malformed reminder can leave the normal path after repeated failure instead of consuming every polling cycle. Redrive then becomes an explicit recovery action after the data or downstream condition has been corrected.
Duplicates count.
Consider one EU renewal due at 09:00 local business time. The scheduler publishes the event, worker A sends it, and the process stops before acknowledgement. Worker B later receives the same event. A polling table has exactly the same dangerous interval if its process sends before marking the row complete; changing the transport doesn't erase the gap. The queue design is still easier to reason about because the repeat delivery is explicit, the message can move away from healthy traffic after repeated rejection, and an operator has a named redrive action. The database remains essential, but its job narrows to storing the deadline, the stable business key, claim state, and audit result. This division also makes the eval precise: deliver that event twice, interrupt it at the send/ack boundary, and require one business effect plus a visible transport outcome.
Database polling still has a place. It's easy to inspect with SQL, and a single database can be an acceptable dependency for a genuinely small system. The catch is that exponential backoff, concurrency control, claim expiry, and stuck-job visibility become application features that somebody must design, test, and operate. Don't call the poller cheaper unless that engineering work is included.
Python implementation of the business idempotency boundary
Standard queues are at-least-once, so a worker must assume that the same renewal event can arrive again. FIFO deduplication does not remove that obligation: its deduplication window is five minutes. A stable business key such as renewal-reminder:{account_id}:{renewal_date} should cross every retry, and the database should atomically claim it before the irreversible send.
This is the notebook-to-prod test I care about: feed the handler the same event twice and assert that only one send is claimed. No vendor-specific payload is needed to verify that invariant.
import sqlite3
from dataclasses import dataclass
@dataclass(frozen=True)
class RenewalReminder:
account_id: str
renewal_date: str
region: str
@property
def idempotency_key(self) -> str:
return f"renewal-reminder:{self.account_id}:{self.renewal_date}"
def claim_once(connection: sqlite3.Connection, event: RenewalReminder) -> bool:
connection.execute(
"CREATE TABLE IF NOT EXISTS reminder_claims "
"(idempotency_key TEXT PRIMARY KEY, region TEXT NOT NULL)"
)
cursor = connection.execute(
"INSERT OR IGNORE INTO reminder_claims (idempotency_key, region) VALUES (?, ?)",
(event.idempotency_key, event.region),
)
connection.commit()
return cursor.rowcount == 1
def main() -> None:
connection = sqlite3.connect(":memory:")
event = RenewalReminder(
account_id="acct_1042",
renewal_date="2026-09-01",
region="eu",
)
assert claim_once(connection, event) is True
assert claim_once(connection, event) is False
print("duplicate delivery suppressed")
if __name__ == "__main__":
main()
The example deliberately stops at the claim boundary. In a real sender, the database transaction and the external side effect cannot usually be committed atomically, so the delivery provider should also receive the stable key where it supports idempotency. A 429 is retryable only after backoff, honoring Retry-After when present; a blind tight loop just converts throttling into pressure. The exact send contract belongs in an integration test and eval harness, including a duplicate event, a delayed retry, and a crash at each state transition.
Keep a durable application record even after success. Queue retention is finite, and acknowledgement removes a message, so the queue is transport rather than the audit ledger.
Audit state stays put.
Regional deadline data, cron limits, and audit governance
A cron trigger can enqueue reminders that have crossed their business deadline. It should not spend minutes delivering a large batch itself. Infrai cron executions, for example, are limited to 900 seconds, accept only a public http_url, do not backfill triggers missed while paused, and can have second-level timing jitter. The safe shape is cron-to-queue-to-worker.
For region-specific deadlines, store the intended instant and its source timezone rather than inferring either from worker location. I'm not sure which legal or product rule defines "business deadline" for every US and EU account; that needs a product decision and tests around daylight-saving transitions. Infrastructure cannot settle that ambiguity.
Long processing should always be handed to workers. Short trigger. Clear ownership.
Provider comparison through the retry boundary
| Option | Retry and failure isolation | Operational fit | Choose it when | Avoid it when |
|---|---|---|---|---|
| Manual database poller | Application-owned backoff, claims, and failure table | Few moving parts at first | Volume is tiny and the team accepts owning queue semantics | Poison-message isolation and operator redrive already matter |
| AWS SQS | Queue delivery with FIFO documentation available | Managed cloud component | The application is already centered on AWS | Cross-provider coupling is a primary constraint |
| Celery | Python worker ecosystem | Worker-oriented application stack | Python task execution belongs inside the app deployment | A plain language-neutral HTTP boundary is required |
| Temporal | Workflow orchestration model | Long-running workflow platform | The job is really a multi-step durable workflow | One delayed reminder would not justify workflow machinery |
| Infrai | Queue, acknowledgement, DLQ, and redrive capabilities behind REST | One key and one billing relationship across backend capabilities | A self-describing HTTP API is preferable to another SDK | DAGs, fan-out joins, or Kafka-style replay are required |
Those choices aren't interchangeable. Temporal addresses an orchestration class that a basic queue does not. Celery is a natural candidate when Python workers are already the unit of deployment. AWS SQS deserves preference when AWS-native operation and its FIFO model align with the rest of the stack. The database poller wins when its intentionally narrow scope is a feature, not an optimistic forecast.
Infrai is worth considering when the team wants queue capabilities without installing another SDK. Its public discovery surface describes each capability with request and response schemas, billing metadata, and runnable examples; reading the capability endpoint is the integration step. That is a concrete advantage for a Python AI app that may later need other backend capabilities through the same REST conventions and key.
I wouldn't choose it for a workflow that needs DAG orchestration or a fan-out/fan-in join. Delayed messages top out at seven days, bodies at 256 KB, and retention at 30 days; acknowledged messages are deleted, with no Kafka-style replay or multiple consumer groups. Push subscriptions also require a public HTTPS target. Those are design boundaries, not footnotes. Stick with Temporal for durable multi-step workflows, and evaluate Kafka when replay and independent consumer groups define the problem.
The API's idempotency convention is useful here: documented idempotent capabilities accept an Idempotency-Key, have a deterministic server-derived fallback, and use a 24-hour default deduplication window. Still, application idempotency remains mandatory for a standard queue. Transport deduplication and business deduplication protect different boundaries.
The following acknowledgement helper is intentionally narrow. It uses the verified verb-style route, reads both the base URL and key from the environment, declares the method, surfaces 4xx responses, and backs off on 429. Supply the acknowledgement body exactly as returned by live discovery rather than guessing its fields.
import os
import time
import requests
def acknowledge(payload: dict[str, object]) -> dict[str, object]:
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
url = f"{base_url}/v1/queue/ack"
for attempt in range(5):
response = requests.request(
method="POST",
url=url,
headers={"Authorization": f"Bearer {api_key}"},
json=payload,
timeout=30,
)
if response.status_code != 429:
response.raise_for_status()
return response.json()
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("acknowledgement remained rate-limited after five attempts")
Failure evaluation before rollout
Measure the age of the oldest ready message, retry count by error class, DLQ depth, time from deadline to successful delivery, duplicate claims rejected, and redrive outcomes. These are evals for infrastructure: they turn "retry works" into assertions that can fail before customers discover the gap.
Run a small failure matrix before choosing a provider. Inject a 429 with Retry-After, terminate a worker after it claims an event, deliver the same event twice, place one malformed message ahead of valid work, and pause the scheduler across a deadline. Your mileage may vary with traffic shape, but those cases expose the semantic difference between a queue and a polling loop faster than a feature checklist.
The decision rule is compact. Choose queue plus DLQ and redrive when failed work must be visible and recoverable; retain the application database for business state and audit history. Keep manual polling only when the team can name the limited scale, the ownership model, and the point at which it will migrate.
Top comments (0)