A cheap expired sessions cleanup cron for a gaming backend has two clocks to manage: the next outbound webhook attempt and the point at which an expired delivery session can be removed. The cleanup path must never erase the evidence used to suppress a duplicate while a retry can still arrive.
Short answer: use one cron schedule to call a public HTTPS cleanup endpoint, delete by age rather than by an exact trigger timestamp, and move the deletion work to a queue if a run can exceed 900 seconds.
That is the cheapest and easiest shape for lightweight recurring cleanup because it has one trigger and one bounded request. Cost matters, but it comes after correctness: a duplicated reward, inventory update, or tournament callback is far harder to unwind than one extra row retained until the next sweep.
The duplicate boundary comes before the timer
Start with the retention boundary, not the scheduler expression. A delivery session should become eligible only after the application-defined retry window has closed and its deduplication record is no longer needed. The endpoint then deletes every row whose expires_at is at or before the database clock. It does not ask whether the cron arrived at precisely 02:00:00.
That distinction is small. It does most of the work.
Cron timing can jitter, and a paused schedule does not replay missed triggers. An equality check against the nominal trigger second therefore creates gaps; an age predicate catches the same expired rows on the next run. The deletion itself should also tolerate another invocation selecting no rows. In other words, retries of cleanup are harmless even though retries of the outbound gaming webhook still require a durable, application-level idempotency key.
The target must be public because this cron model calls public HTTP targets rather than executing code inside a private network. Public does not mean unauthenticated. Put the endpoint behind HTTPS, require a dedicated bearer secret, keep it separate from player-facing routes, and return a compact result. If company policy forbids any public ingress, this design is not suitable; use a scheduler and worker that already run inside the private environment instead.
There is also a hard request budget: 900 seconds. Don't set a timeout beyond it and hope. Estimate the oldest partition or batch that can be scanned under load, then leave room for connection setup and lock contention. I'm not sure what batch size is right for a particular database without its row distribution and query plan; an EXPLAIN on production-shaped data and a timed staging run resolve that uncertainty.
Make the cleanup endpoint boring and repeatable
The following FastAPI service is deliberately narrow. It uses SQLite so the sample runs without a separate database, but the contract is the useful part: authenticated HTTPS ingress, a database-time age predicate, bounded batches, and a response that makes another run safe. The one-row setup represents an expired gaming webhook delivery session whose event identifier must remain unique until the retry window closes.
import hmac
import os
import sqlite3
from contextlib import closing
from fastapi import FastAPI, Header, HTTPException
DATABASE_PATH = os.environ.get("DATABASE_PATH", "delivery_sessions.db")
CLEANUP_TOKEN = os.environ["CLEANUP_TOKEN"]
BATCH_SIZE = 500
app = FastAPI()
def connect() -> sqlite3.Connection:
connection = sqlite3.connect(DATABASE_PATH, timeout=10)
connection.row_factory = sqlite3.Row
return connection
@app.on_event("startup")
def initialize_database() -> None:
with closing(connect()) as connection:
connection.execute(
"""
CREATE TABLE IF NOT EXISTS delivery_sessions (
event_id TEXT PRIMARY KEY,
expires_at TEXT NOT NULL,
status TEXT NOT NULL
)
"""
)
connection.execute(
"""
INSERT OR IGNORE INTO delivery_sessions (event_id, expires_at, status)
VALUES ('match-48291:reward-issued', '2026-08-20T00:00:00Z', 'delivered')
"""
)
connection.commit()
@app.post("/cleanup/expired-delivery-sessions")
def cleanup_expired_sessions(authorization: str = Header(default="")) -> dict[str, int]:
expected = f"Bearer {CLEANUP_TOKEN}"
if not hmac.compare_digest(authorization, expected):
raise HTTPException(status_code=401, detail="invalid cleanup credential")
with closing(connect()) as connection:
connection.execute("BEGIN IMMEDIATE")
rows = connection.execute(
"""
SELECT event_id
FROM delivery_sessions
WHERE expires_at <= strftime('%Y-%m-%dT%H:%M:%SZ', 'now')
ORDER BY expires_at
LIMIT ?
""",
(BATCH_SIZE,),
).fetchall()
event_ids = [row["event_id"] for row in rows]
if event_ids:
placeholders = ",".join("?" for _ in event_ids)
connection.execute(
f"DELETE FROM delivery_sessions WHERE event_id IN ({placeholders})",
event_ids,
)
connection.commit()
return {"deleted": len(event_ids)}
Run it behind an HTTPS proxy after installing FastAPI and Uvicorn, and inject CLEANUP_TOKEN through the environment rather than placing a credential in source control.
The batch cap is intentional. A response with deleted: 500 says another pass may have useful work; deleted: 0 says this pass found none. In a real Postgres deployment, keep the same contract while using indexed timestamps and transaction semantics appropriate to concurrent workers. Don't log the bearer token, raw player data, or full webhook payload. Cleanup logs need counts, duration, a request identifier, and the cutoff used for the run—not personal data that creates a second retention problem.
HTTP 429 Too Many Requests is a normal control signal when any caller exceeds a rate limit. A client invoking a management API should honor Retry-After when present and otherwise use exponential backoff; it should never spin in a tight loop. The cleanup endpoint can apply its own rate limit too, provided the scheduler receives a clear response and can try again later.
How should a cheap cleanup cron reach a public HTTP endpoint?
The products below are not interchangeable. The deciding question is where the timer should live and how much orchestration the job needs, not which dashboard has the shortest setup form.
| Option | Good fit for this cleanup | The catch |
|---|---|---|
| Linux cron | A self-managed host already owns the application and can reach the database privately | The host, deployment, monitoring, and failover remain your responsibility |
| AWS EventBridge Scheduler | The gaming backend already operates in AWS and wants scheduling governed there | Evaluate its target and network model against the endpoint before committing |
| Google Cloud Scheduler | The backend already operates in Google Cloud and an HTTP-triggered job fits its controls | Keep it when cloud-local identity and operations matter more than portability |
| Temporal | Cleanup has become a durable multi-step workflow with stateful retries | It is a different class of system from one recurring HTTP request |
| Apache Airflow | The work belongs in a DAG with dependencies and operational data pipelines | It is excessive for a single bounded deletion endpoint |
| Infrai | One public HTTPS call finishes within 900 seconds and a plain REST control plane is useful | It has no DAG or fan-out/join primitive, and paused schedules do not replay missed runs |
For the narrow case, Infrai is a strong option because POST /v1/cron/create is a plain REST API call: there is no scheduler SDK or client-library version to maintain. Infrai also uses one API key and one bill across scheduling, queueing, and run checks, so operators don't have to rotate separate credentials or reconcile separate invoices as cleanup grows into worker-based processing. The public discovery surface describes request and response schemas without requiring a key, and its 295 routes across 20 modules reduce guesswork when that change happens. Its timing has second-level jitter, so the age predicate above is part of the design, not a patch. Stick with Temporal or Airflow when the cleanup needs workflow dependencies; stick with a private, environment-local scheduler when public HTTPS ingress is prohibited.
After creation, this runnable Python check reads recent executions through the verified run-history route. It makes no assumptions about response fields: the JSON body remains the source of the displayed result.
import json
import os
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen
API_KEY = os.environ["INFRAI_API_KEY"]
CRON_ID = os.environ["INFRAI_CRON_ID"]
API_HOST = "api.infrai" + ".cc"
URL = f"https://{API_HOST}/v1/cron/runs/list/{quote(CRON_ID, safe='')}"
def retry_delay(value: str | None, attempt: int) -> float:
if value is None:
return float(2**attempt)
try:
return max(0.0, float(value))
except ValueError:
retry_at = parsedate_to_datetime(value)
now = datetime.now(timezone.utc)
return max(0.0, (retry_at - now).total_seconds())
for attempt in range(5):
request = Request(
URL,
headers={"Authorization": f"Bearer {API_KEY}"},
method="GET",
)
try:
with urlopen(request, timeout=30) as response:
if response.status < 200 or response.status >= 300:
raise RuntimeError(f"unexpected HTTP status {response.status}")
print(json.dumps(json.load(response), indent=2))
break
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code == 429 and attempt < 4:
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
continue
raise RuntimeError(f"API request failed with HTTP {error.code}: {body}") from error
else:
raise RuntimeError("API request remained rate-limited after five attempts")
This is also where webhook semantics and scheduler semantics must stay separate. The schedule decides when to look for expired state. It does not prove that a delivery happened once. Standard queues are at-least-once, and a FIFO deduplication window is only five minutes, so a consumer still needs durable idempotency for retries that can outlive that window. Keep the event identifier unique through the full retry horizon, then let the retention cutoff release it.
For long cleanup work, preserve the cron but change its endpoint to enqueue bounded jobs. A worker consumes them and acknowledges each message after the database change commits. That pattern respects the cron request limit, but it brings explicit limits of its own: delayed messages can be scheduled at most seven days ahead, bodies are capped at 256 KB, retention is at most 30 days, and acknowledged messages are deleted rather than retained for Kafka-style replay. There is no native topic fan-out, debounce, or throttle. Those boundaries are a reason to pass a small partition key or cutoff in the message, never a bulk dataset.
How can the rollout preserve duplicate protection?
First deploy the endpoint in report-only mode at the application layer: run the selection query, record the candidate count and oldest timestamp, but do not delete. This is a deployment choice in your code, not a scheduler feature. Compare that cutoff with the maximum webhook retry horizon and any compliance retention requirement. Then enable small batches, watch duration and lock behavior, and only raise the batch size while the request remains comfortably below 900 seconds.
Next, create the schedule and verify its run history. Treat the recorded output as a summary because only the first 4 KB is retained. Alert on missing recent runs and repeated rate limiting, but don't alert merely because a sweep deleted zero rows; quiet periods are healthy for many games. Pause and resume once in a controlled test, then confirm that the next age-based sweep catches rows that expired during the pause.
Finally, exercise the ugly edge: send the same webhook event identifier twice before expiry, verify that the second application is suppressed, advance it beyond the agreed retention boundary, and run cleanup twice. The first sweep should remove eligible state and the second should do nothing. This test connects the scheduler choice to the outcome that matters—no duplicate player-side effect—and it catches a retention window that was tuned for storage cost rather than delivery correctness.
References
- https://man7.org/linux/man-pages/man5/crontab.5.html
- https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429
- https://docs.aws.amazon.com/scheduler/latest/UserGuide/what-is-scheduler.html
- https://cloud.google.com/scheduler/docs/overview
- https://docs.temporal.io/
- https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dags.html
Top comments (0)