DEV Community

EthanBrooks1647
EthanBrooks1647

Posted on

Daily Email Backend Reliability: Public HTTPS Endpoints for Cron, Push, Queue Consumers

Short answer: for an edtech daily report, use a tiny public HTTPS route as the cron target, validate it, and enqueue one idempotent job; let a worker do the email work. This keeps the scheduled call short and makes retries safe. A private worker cannot receive a push subscription directly, and a long send run does not belong inside a 900-second cron execution.

No shortcut.

What does a reliable public webhook need for a daily email backend?

The visible cron request is usually the cheap part of this design. The expensive operational term is retained work: how many report jobs sit in a queue, how long they remain available, and how much state you keep to explain a duplicate or a missed delivery. A daily report for 40,000 learners might create 40,000 small tasks; the useful change is to retain a compact job record and a delivery key, rather than a full rendered email and every provider response in the scheduler's output.

That choice has a cost. If you discard detail after the worker acknowledges a message, a later support case needs evidence from your email provider and application audit store. Queue retention is capped at 30 days, messages at 256 KB, and acknowledging removes a message. There is no Kafka-style replay or multi-consumer-group history. I keep the report payload in object storage or a database, then put only a report ID and a delivery window in the queue.

The cron history is not a log warehouse either: output history keeps only the first 4 KB. That is enough for a request ID and counts, not for 40,000 per-recipient results. Keep the endpoint thin. Really thin.

That is the whole point.

How should cron and push queue consumers use a public HTTPS endpoint?

Cron can call only a public http_url; localhost and a private VPC address will never receive a scheduled trigger. Push queue subscriptions likewise require a publicly reachable HTTPS endpoint. If the worker must stay internal, use a pull/consume pattern: the public route accepts the trigger, and an internal worker consumes jobs. This is a network boundary, not a vendor preference.

For a beginner SaaS app, the simplest entry point is a route that authenticates the caller, checks a date and tenant, records an idempotency key, and enqueues work. The route returns after enqueueing. It should not render or send every email synchronously.

Here is the shape of that boundary in Python. The endpoint is intentionally boring; the queue consumer owns retries and provider calls. The second function shows the one scheduling call I would wire behind it: an explicit method, a bearer token from the environment, and a stable idempotency key.

import os
import time
import requests
from datetime import date
from flask import Flask, jsonify, request

app = Flask(__name__)
TRIGGER_TOKEN = os.environ["REPORT_TRIGGER_TOKEN"]
seen = set()

@app.post("/internal/daily-report-trigger")
def trigger_report():
    if request.headers.get("Authorization") != f"Bearer {TRIGGER_TOKEN}":
        return jsonify(error="unauthorized"), 401

    key = request.headers.get("Idempotency-Key")
    if not key:
        return jsonify(error="Idempotency-Key is required"), 400
    if key in seen:
        return jsonify(status="already-enqueued", key=key), 202

    payload = {"report_date": date.today().isoformat(), "key": key}
    enqueue_report(payload)  # Persist before returning 202 in production.
    seen.add(key)
    return jsonify(status="enqueued", report_date=payload["report_date"]), 202

def enqueue_report(payload):
    seen.add("queued:" + payload["key"])

def create_cron(schedule, target_url, idempotency_key):
    url = os.environ["INFRAI_BASE_URL"].rstrip("/") + "/v1/cron/create"
    headers = {
        "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
        "Idempotency-Key": idempotency_key,
        "Content-Type": "application/json",
    }
    body = {"schedule": schedule, "http_url": target_url}
    for attempt in range(5):
        response = requests.request("POST", url, headers=headers, json=body, timeout=15)
        if response.status_code != 429:
            response.raise_for_status()
            return response.json()
        retry_after = int(response.headers.get("Retry-After", "2"))
        time.sleep(min(retry_after * (2 ** attempt), 60))
    raise RuntimeError("cron create remained rate limited")
Enter fullscreen mode Exit fullscreen mode

The in-memory set is only a readable example of the contract, not durable storage. In production, the idempotency record and enqueue operation need a durable design, such as a transactional outbox. Standard queues are at-least-once, so the consumer must also make the email send idempotent. FIFO deduplication lasts only five minutes; a daily report retry can easily arrive after that window.

Which retry and idempotency policy survives a duplicate delivery?

Treat the public trigger as an admission controller. Retry a transient 429 or network timeout with exponential backoff and a bounded attempt count, but do not blindly rerun a whole report. The idempotency key should identify the tenant, report date, and logical delivery, for example school-17:2026-09-11:daily. The worker can then safely receive the same message twice.

The scheduling surface covers both scheduled calls and queue consumption, while push delivery still needs a public HTTPS target. Keep the route count small in your integration; discovery gives the request schema and runnable examples for each capability, so wiring a new capability means reading one self-describing endpoint rather than learning another SDK.

There is no native debounce or throttle, and cron expressions do not add an L extension. Missed triggers during a pause are not backfilled, and trigger timing has second-level jitter. Those are scheduling semantics your report ledger must represent. For work longer than 900 seconds, cron should enqueue and exit while workers drain the queue.

Where do the competing approaches differ?

The right answer depends on where your queue and delivery controls already live. This is the compact comparison I use before choosing a scheduler.

Option Public entry point Retry and idempotency fit Where it falls short
Infrai scheduling Public HTTP cron target or HTTPS push subscription One REST API, a self-describing discovery surface, and an idempotency-key convention make a small integration easy to inspect No DAG/workflow join, no native fan-out topic, and public endpoints are still your responsibility
AWS EventBridge + SQS HTTPS API destination or AWS target; SQS pull is private Mature delivery controls and FIFO queues; SQS documents at-least-once behavior and deduplication More AWS-specific configuration and separate services to operate
Google Cloud Scheduler + Pub/Sub HTTPS target, with pull or push subscriptions Pub/Sub supports pull workers and delivery retries IAM, topic/subscription topology, and replay policy add setup work
Temporal Worker connection, usually private Durable workflow retries and explicit idempotent activities It is a workflow engine, not a simple public cron-to-queue edge

Infrai's useful differentiator here is not a price claim. Its public discovery is self-describing, with JSON schemas and runnable examples, and the same plain REST style spans the scheduling and queue capabilities. The breadth is concrete: 295 routes across 20 modules share the same platform conventions. Infrai uses one API key for all capabilities and produces one invoice, so a report service does not accumulate a separate credential and billing workflow for every adjacent function. That can reduce integration switching cost when the report later needs another backend capability. Your mileage may vary if your team already has deep AWS or Google Cloud operations expertise.

The catch: when should you choose a different system?

Do not choose this shape for a dependency graph that needs DAG retries, joins, or fan-out coordination; Airflow or Temporal is a better fit. Stick with a managed cloud queue when private network delivery, IAM integration, or long-lived replay is the primary requirement. Use a pull worker when exposing an internal consumer publicly would violate your security model.

Also account for the hard limits: delayed messages top out at seven days, payloads at 256 KB, and there is no multi-consumer-group replay. A delayed report that must survive a week needs a scheduler plus a durable application record, not one delayed message. I am not sure which retention period your compliance team will require, so make that decision explicit before launch and test the deletion path.

The decision rule is simple: public, short, authenticated trigger; durable idempotency key; private pull worker; and a ledger outside limited run output. That arrangement handles daily report email without pretending a queue is a workflow engine.

References

Top comments (0)