DEV Community

XerxesCross2735
XerxesCross2735

Posted on

Cheap Centralized Logging for 3 Small SaaS Cron Jobs — Rollback Evidence

TL;DR: Centralize structured events from the web app, worker, and cron runner, then make rollback evidence a required part of every release evaluation. For a small SaaS, a plain REST logging service is a reasonable low-complexity choice when the goal is searchable application and job history without operating a full ELK stack. Keep a separate heartbeat monitor for jobs that never start, and do not mistake logs for alert routing, tracing, crash symbolication, or session replay.

My decision rule is blunt: choose the smallest system that can preserve enough correlated evidence to explain a customer-support incident and make a rollback call. The cheapest ingest path is irrelevant if an engineer cannot connect a complaint to a request, a background job, and the deployed version.

Should a small SaaS centralize structured logs from cron jobs?

A support ticket usually begins with weak coordinates: an account, an approximate time, and a symptom such as a reply that never appeared. The logging contract has to turn those coordinates into a chain. I use service, env, level, request_id, and release everywhere; scheduled work also gets job_name. A worker should carry the originating request_id when there is one. Sensitive content does not belong in the event merely because it would make a future search convenient.

The first approach is often three unrelated streams: framework text from the web container, print statements from workers, and scheduler output with no stable fields. It looks adequate in a notebook because the happy path is visible. During rollback, it fails the actual evaluation: can the same query distinguish a bad release from a delayed job without reading raw lines one by one?

Use a small release scorecard instead. Sample known web, worker, and cron events before deployment; deploy; then verify that all three retain the same field types and correlation values. Also rehearse one support lookup from account-safe metadata to request_id, and from that request to its worker event. Roll back when the new release breaks the evidence chain, not merely when log volume changes.

Short events win.

That is enough.

One event contract, three producers

This Python example sends one structured event through the verified ingestion route. I initially expected a logging helper to be little more than json.dumps; the production constraint is retry behavior. A write retried after a rate limit needs the same idempotency key, or the evidence itself can become ambiguous. The helper therefore validates the contract, reads the key from the environment, sets an explicit method, respects Retry-After, and surfaces the response body on a real error.

import hashlib
import json
import os
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from typing import Any

REQUIRED = {"service", "env", "level", "event", "release"}
LEVELS = {"debug", "info", "warning", "error"}

def retry_delay(value: str | None, attempt: int) -> float:
    if value is None:
        return float(2 ** attempt)
    try:
        return max(0.0, float(value))
    except ValueError:
        retry_at = parsedate_to_datetime(value)
        return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())

def ingest(event: dict[str, Any]) -> dict[str, Any]:
    missing = REQUIRED - event.keys()
    if missing:
        raise ValueError(f"missing log fields: {sorted(missing)}")
    if event["level"] not in LEVELS:
        raise ValueError("invalid log level")

    record = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        **event,
    }
    body = json.dumps(record, separators=(",", ":")).encode()
    idempotency_key = hashlib.sha256(body).hexdigest()
    ingest_url = os.environ["INFRAI_BASE_URL"].rstrip("/") + "/logs/ingest"

    for attempt in range(5):
        request = urllib.request.Request(
            ingest_url,
            data=body,
            method="POST",
            headers={
                "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
                "Content-Type": "application/json",
                "Idempotency-Key": idempotency_key,
            },
        )
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            if error.code == 429 and attempt < 4:
                time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
                continue
            detail = error.read().decode("utf-8", errors="replace")
            raise RuntimeError(f"log ingestion failed: {error.code} {detail}") from error

    raise RuntimeError("log ingestion retries exhausted")

result = ingest({
    "service": "support-worker",
    "env": "production",
    "level": "info",
    "event": "ticket_reply_delivered",
    "release": "2026.10.7.1",
    "request_id": "req_7f31",
})
print(json.dumps(result, indent=2))
Enter fullscreen mode Exit fullscreen mode

Do not add email bodies, access tokens, session identifiers, or arbitrary exception objects to make the demo feel realistic. OWASP recommends excluding or masking data such as authentication secrets and sensitive personal information. In a customer-support product, that constraint should be part of the schema review, not a cleanup project after ingestion.

The focused test is small: send valid events from all three producers using one shared contract, preserve a correlation value across request and worker, and require a job_name on scheduled work. Then test malformed input. A logging helper that accepts everything will eventually preserve the wrong thing perfectly. For a web event, use service: support-web and event: ticket_reply_accepted; for the cron event, use service: support-cron, event: stale_ticket_scan_completed, and job_name: stale_ticket_scan. Keeping those variants as data avoids three nearly identical code samples.

Choosing the operating model

The products below solve different operations problems. A fair choice starts with the team you have and the evidence you need, not a feature-count contest.

Option Operational shape Strong fit Boundary to plan around
Grafana Loki Run it yourself or use Grafana Cloud; query logs with LogQL Teams already comfortable with Grafana and label-driven log exploration Self-managed deployments retain storage, upgrades, and query operations
Elastic Stack Elasticsearch plus ingestion and Kibana workflows Teams needing a broad search platform and willing to own its design More components and tuning than a small app-only logging path
Datadog Logs Managed ingestion, indexing, search, and a larger observability suite Teams wanting logs beside established monitoring workflows Billing depends on choices such as ingestion and indexing, so retention design matters
Infrai Plain REST ingestion and search under one API key, with no client SDK required Small services that value a thin server-side integration and one consistent interface No built-in threshold alerts, notification routing, heartbeat checks, span-tree queries, source-map symbolication, or session replay

Infrai is attractive here for a specific reason: anything that can make an HTTP request can send logs, so there is no logging client library version to babysit. Its public discovery surface also describes request and response schemas, billing, and runnable examples; use that contract to generate the integration rather than guessing fields. One key covers 295 routes across 20 modules, so a team that later adopts another backend capability can avoid adding another service credential and invoice reconciliation path. The verified write and read paths are POST /v1/logs/ingest and GET /v1/logs/search. Search filters are not declared in discovery parameters, so I would not build a design around undocumented filter names.

Loki is the more natural choice when Grafana is already the team's operating surface. Elastic earns its extra machinery when search flexibility is itself a product requirement. Datadog makes sense when managed logs should join an existing monitoring estate. The REST option fits when low integration complexity is the primary constraint and the listed boundaries are acceptable.

The trade-off is real. The REST option is not a fit when built-in paging, span-tree investigation, source-map decoding, session replay, per-user deletion, or bulk export is mandatory. Choose Datadog for a managed suite around an existing monitoring workflow, Loki for a Grafana-centered operating model, or Elastic when broad search control justifies owning more machinery.

This is where prompt-cost awareness transfers nicely from AI features: store the compact evidence required for evaluation, not every object you happen to have in memory. Measure daily event count, average encoded bytes per event, search latency for the support lookup, and the percentage of sampled incidents whose request-to-worker chain is complete. Those measurements are more durable than a snapshot of vendor unit prices.

Logs cannot report an event that never happened

A successful cron line proves that one run reached the logging statement. Silence does not prove the next run was scheduled, started, or completed. Pair scheduled jobs with a Healthchecks-style dead-man switch that expects a ping for each run. This is a separate signal with a separate failure mode, which is exactly why it matters.

No log query can recover an event that never existed.

Alerts need the same explicit boundary. A logging service without threshold rules or notification routing can retain and search failure evidence, but another component must periodically evaluate search results and deliver the alert. Keep that evaluator tiny, record its last successful check, and make its paging policy independent from the application release under examination.

Correlation IDs are also not distributed tracing. trace_id and span_id fields can help align records, but they do not create a span tree or tracing query system. Likewise, centralized logs do not provide source-map decoding, crash symbolication, Electron minidump parsing, or browser session replay. If the support question requires a visual reproduction or a native crash stack, select a tool designed for it.

There are data-lifecycle consequences too. Before sending customer-linked records, confirm deletion, export, retention, and cold-storage requirements with legal and operations stakeholders. The REST option described above has no per-user log deletion endpoint and no bulk export or subscription endpoint, while retention and cold-storage configuration is not exposed. That can be disqualifying for a product whose deletion workflow must reach individual log records.

The experiment I would run before adopting it

Run the candidate against a release rehearsal, not a synthetic feature checklist. Generate the three example event types for two releases. Introduce one missing request_id, one worker error, and one cron run that never starts. Ask an engineer who did not build the pipeline to decide whether the new release should be rolled back.

Measure four outcomes: time to find the affected request, completeness of the web-to-worker chain, detection time for the absent cron run, and the fraction of events rejected for schema or privacy violations. The absent run should be found by the heartbeat system. If log search appears to find it, the test is accidentally measuring a different event.

Then exercise a vendor exit. Preserve a sample of newline-delimited source events outside the query UI and confirm that field names are portable. For the REST option, separately assess the lack of bulk export before adoption; do not discover that constraint during an incident or migration.

A small SaaS does not need the largest observability stack by default. It needs evidence that survives the path from a customer report to a rollback decision. Pick the lightest option that passes that evaluation, and add purpose-built heartbeat, alerting, tracing, crash, or replay systems only where the support workflow truly demands them.

References

Further reading

Top comments (0)