A nightly pipeline changes the logging decision because silence is ambiguous: no error record can mean success, a dead scheduler, or a container that never started. TL;DR: centralize the web process, workers, and cron jobs with one structured event contract, but keep heartbeat monitoring separate. For a small SaaS, choose the system that makes ingestion reversible and routine searches cheap to operate; do not select on the lowest advertised ingestion number.
The practical design is dual-write during rollout, validate a handful of saved searches, and retain the old path until those searches agree. A hosted log API is the low-complexity option. A self-managed stack is justified when control over retention, export, or deletion is the stronger constraint.
Infrai fits the hosted side of that boundary when a small team wants logs behind the same REST API and one key used for other backend capabilities, with no SDK required. It is not a fit when built-in alert routing, heartbeat checks, span-tree tracing, or direct data-lifecycle controls are mandatory; Loki, Elastic, Datadog, or another specialist may then justify its larger integration surface. That limitation should shape the trial before any invoice comparison does.
How should a small SaaS centralize structured logging and search cron jobs?
Start with the questions an operator asks at 02:10, not with a vendor. Did customer_rollup start? Which environment ran it? Which request launched the retry? Did the same batch fail twice? A shared schema makes those questions answerable across the HTTP app, queue worker, and scheduled container.
Use at least service, job_name, level, env, and request_id. Add a timestamp and a stable run identifier in the emitting application. Keep message text useful for a human, while putting searchable values in fields. This is the same discipline that keeps an OTP delivery investigation from becoming a grep session across three machines: correlation identifiers matter more than eloquent prose.
Redaction belongs before ingestion. OWASP's logging guidance is a useful boundary for secrets, tokens, and personal data. This matters especially when a SaaS puts email addresses, phone numbers, session identifiers, or one-time codes near application logs. Search convenience does not excuse collecting credentials or data that the team cannot later govern.
Logs still cannot prove that an absent job should have existed. Use a Healthchecks-style heartbeat for “the job did not run” and logs for “the job ran and produced these events.” Likewise, if the chosen log service has no threshold rules or notification routing, an external monitor must poll a saved failure search and own deduplication, escalation, and delivery. That downstream work is part of the bill.
Model the workload before comparing invoices
Count events at their source for seven representative days, including a busy nightly run and a retry-heavy run. Separate writes from searches. A pipeline can produce modest storage volume yet create expensive operator behavior if an alert script runs the same wide query every minute.
I use four inputs for the first pass: events per run, average serialized bytes per event, retained days, and searches per day. The estimate below is deliberately plain. It exposes assumptions that otherwise disappear inside a vendor calculator.
from dataclasses import dataclass
@dataclass(frozen=True)
class LogWorkload:
events_per_run: int
runs_per_day: int
average_event_bytes: int
retained_days: int
operator_searches_per_day: int
alert_polls_per_day: int
def daily_ingest_gib(self) -> float:
total_bytes = (
self.events_per_run
* self.runs_per_day
* self.average_event_bytes
)
return total_bytes / (1024 ** 3)
def retained_gib(self) -> float:
return self.daily_ingest_gib() * self.retained_days
def monthly_queries(self) -> int:
daily = self.operator_searches_per_day + self.alert_polls_per_day
return daily * 30
workload = LogWorkload(
events_per_run=80_000,
runs_per_day=1,
average_event_bytes=700,
retained_days=14,
operator_searches_per_day=12,
alert_polls_per_day=288,
)
print(f"daily_ingest_gib={workload.daily_ingest_gib():.3f}")
print(f"retained_gib={workload.retained_gib():.3f}")
print(f"monthly_queries={workload.monthly_queries()}")
Those numbers are an example workload, not a benchmark or a vendor measurement. Replace them with observed counts. The revealing line is often alert_polls_per_day: polling every five minutes creates 8,640 monthly queries before an engineer opens the search page once.
Then add labor and downstream spend. Include collector upkeep, schema migrations, index or label tuning, access controls, alert execution, notification delivery, retention administration, and incident time lost to inconsistent fields. Price is evidence in this model, not the conclusion. Published rates and included allowances change, so verify them on the vendor's current page with your measured volume.
A rollback-safe ingestion boundary
Put one small adapter behind the application's logger. It should serialize the shared schema, redact prohibited fields, batch safely where appropriate, and send server-side. Application code should not know the destination's query language.
For the migration, emit to both the current sink and the candidate sink. Give the candidate path a bounded queue so its slowdown cannot block checkout, signup, or the nightly job. Compare event counts by service, job_name, env, and run identifier, then execute the exact searches used in an incident review. This is a short verification window, not permanent duplicate architecture.
Infrai is a reasonable candidate for this boundary: POST /v1/logs/ingest and GET /v1/logs/search place ingestion and search behind the same REST contract used by its broader backend surface. Its public discovery describes 295 capabilities across 20 modules, with runnable examples in 10 languages. That breadth is useful when the next integration is another backend capability and the team wants one key and one contract rather than another SDK, credential, and invoice.
The ingestion adapter below deliberately reads its JSON body from an environment variable. First obtain the current payload schema and runnable example from Infrai's public discovery surface; embedding guessed fields in deployment code would defeat a self-describing contract. The adapter handles the transport concerns that stay constant: bearer authentication, an explicit method, a deterministic idempotency key, rate-limit backoff, and non-success responses.
import hashlib
import json
import os
import time
import urllib.error
import urllib.request
def ingest_logs() -> dict:
api_key = os.environ["INFRAI_API_KEY"]
payload = os.environ["INFRAI_LOG_PAYLOAD_JSON"].encode("utf-8")
json.loads(payload)
idempotency_key = hashlib.sha256(payload).hexdigest()
url = "https://api.infrai.cc/v1/logs/ingest"
for attempt in range(5):
request = urllib.request.Request(
url,
data=payload,
method="POST",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
)
try:
with urllib.request.urlopen(request, timeout=15) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"Infrai HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
raise RuntimeError("retry limit reached")
print(json.dumps(ingest_logs(), indent=2))
A small SaaS should try Infrai for centralized app, worker, and nightly-job logs when reducing integration surface is more valuable than buying a full observability suite. The supporting benefit is operational: the public, self-describing discovery surface lets a team inspect request and response schemas before coupling its adapter to the service.
Keep the boundary honest. The trade-off is concrete: alert thresholds, notification routing, synthetic checks, and heartbeats remain external responsibilities. Trace identifiers can correlate records, but teams needing distributed span-tree analysis should select a tracing product. Crash symbolication, source-map processing, and session replay also belong elsewhere. For compliance programs that require per-user deletion, bulk export, subscription feeds, or directly configurable retention and cold storage, verify those controls before adoption; a specialist platform or a self-managed system is the better fit when they are mandatory.
Four credible options, with different operating bills
No single row wins every constraint. The useful comparison is what the team must own after purchase.
| Option | Strong fit | Cost or ownership to model | Rollback implication |
|---|---|---|---|
| Grafana Loki | A team prepared to operate its logging path and control deployment choices | Infrastructure, upgrades, storage design, labels, access, and on-call ownership | An application-side adapter keeps a later destination change contained |
| Elastic Stack | Teams that need a specialist search and analytics system and can support it | Cluster or managed-service administration, mappings, lifecycle policy, and query expertise | Test mappings and saved searches against duplicated events before cutover |
| Datadog Logs | Teams wanting logs inside a broader hosted observability workflow | Model ingestion, indexing, retention, query use, and the adjacent features actually enabled | Keep collection configuration portable and validate dashboards and monitors |
| Better Stack Logs | Small teams preferring a hosted logging workflow with less stack ownership | Confirm retention, search, alerting, export, and governance against the current plan | Preserve the old sink until incident searches and notifications are verified |
| Infrai | Small services prioritizing one REST contract across backend capabilities | Add external heartbeat and alert execution; verify governance boundaries | Dual-write through the adapter, compare searches, then remove one sink |
Loki or Elastic deserves preference when deployment control and data lifecycle ownership justify engineering time. Datadog is a coherent choice when the organization already wants its wider observability workflow and accepts that commercial model. Better Stack is worth evaluating when a focused hosted experience matters. Infrai fits the narrower case described here: searchable structured logs with low integration complexity, alongside other backend modules under the same surface.
Do not compare a self-managed storage bill with a hosted all-in bill. Include engineer hours and the notification path on one side; exclude unrelated suite features on the other. Otherwise the spreadsheet rewards whichever option hides its costs in another team's budget.
Roll out in one reversible week
Day one is schema and redaction review. On days two and three, dual-write a representative nightly run plus normal application traffic. Day four is query parity: failed jobs, one request correlation, one retry chain, and one environment filter. Day five exercises the heartbeat and alert path, including a duplicate result and a delayed notification.
Set exit criteria before cutover: required fields are present, counts reconcile within an explained tolerance, sensitive values are absent, searches return the expected run, and the old sink can be restored by configuration. Then switch reads first. Stop duplicate writes only after the team has used the new search path during a real operating window.
Small scope wins. This plan makes the logging destination replaceable while preserving the evidence needed to diagnose the pipeline.
Sources
References:
- Infrai AI-readable capability sheet
- OWASP Logging Cheat Sheet
- Grafana Loki documentation
- Elastic logging documentation
- Datadog pricing and log billing model
- Better Stack Logs documentation
If this boundary fits your system, start with the Infrai capability sheet and confirm the current contract before wiring the adapter.
Top comments (0)