Short answer: for a small SaaS notification service, the best inexpensive logging choice is the one that preserves enough evidence to reconstruct a failed delivery without turning every retry into a fresh incident. Start with a compact event schema, separate attempt facts from final outcomes, retain an audit trail keyed by an idempotency key, and compare Datadog, Better Stack, Logtail, Axiom, and self-hosted Loki with the same replayable dataset. Price matters only after the candidates produce equivalent answers. Signal quality comes first.
A payment receipt that fails once, succeeds on retry, and is later logged as both "failed" and "sent" can mislead support, reconciliation, and compliance review. The logging system therefore has two jobs: preserve immutable attempt evidence and expose a derived delivery state without pretending that transport acceptance proves human receipt. Those jobs must survive duplicate workers, delayed callbacks, and partial outages.
Retries lie.
What should cheap app logging prove for a small SaaS?
Begin with questions, not fields. For one logical notification, an operator should be able to determine which template and destination class were involved, how many attempts occurred, what each provider response category was, whether a retry was scheduled, and which terminal state the application recorded. The log must distinguish a duplicate execution from a second business notification.
That requirement implies two identifiers. notification_id names the business action; attempt_id names one dispatch attempt. An idempotency_key links duplicate submissions that must converge on the same business result. None should contain an email address, phone number, message body, access token, or payment data. Store a destination class or an application-generated opaque reference instead. Compliance boundaries vary by jurisdiction and contract, so retention and field classification need explicit review rather than an inherited default.
Use stable event names such as notification.attempt.finished, then put bounded dimensions in structured fields. Prometheus naming guidance is written for metrics, but its insistence that a name represent one logical thing and that labels not create ambiguous dimensions is useful here too. Logs tolerate more detail than metrics; they do not make uncontrolled cardinality free.
package telemetry
import (
"encoding/json"
"io"
"time"
)
type DeliveryAttempt struct {
EventName string `json:"event_name"`
OccurredAt time.Time `json:"occurred_at"`
NotificationID string `json:"notification_id"`
AttemptID string `json:"attempt_id"`
IdempotencyKey string `json:"idempotency_key"`
Channel string `json:"channel"`
Outcome string `json:"outcome"`
FailureClass string `json:"failure_class,omitempty"`
Retryable bool `json:"retryable"`
DurationMS int64 `json:"duration_ms"`
}
func WriteAttempt(w io.Writer, e DeliveryAttempt) error {
return json.NewEncoder(w).Encode(e)
}
Keep outcome and failure_class to reviewed vocabularies. A raw error can help diagnosis, but it belongs in a controlled field with redaction and length limits; otherwise arbitrary provider text becomes both a privacy risk and a grouping key with effectively unbounded cardinality.
Model attempts and outcomes separately
Exactly-once delivery is the wrong claim at a network boundary because a timeout can hide whether the remote service accepted a request. The defensible design is an exactly-once mindset at the application boundary: make submission idempotent, record every attempt, and compute current business state from ordered evidence. Never overwrite the previous attempt because a later one succeeded.
Evidence stays.
Consider three records for one notification: attempt 1 times out after 2,000 ms; attempt 2 receives an acceptance response after 180 ms; a later callback marks the address as rejected. An alert on every non-success record creates two pages for one business action and still risks missing the terminal rejection. An alert on derived state can wait for the retry policy, then notify only when the notification crosses a meaningful boundary such as exhausted or permanent_failure. The attempt records remain searchable for audit and latency analysis.
This is the crucial trade-off: delaying an alert until state is derived reduces noise, but excessive delay can conceal a broad provider failure. Cover that gap with aggregate signals. Count attempts by channel, outcome, and bounded failure class; measure duration; and alert on a sustained change in rates rather than individual retryable events. Do not put unique IDs in metric labels. Their cardinality belongs in logs.
Error grouping needs the same care. Sentry documents that grouping uses event information and permits fingerprints to override grouping behavior. Regardless of backend, a useful fingerprint represents an actionable failure class, not a full message or unique identifier. Group provider_timeout together while leaving IDs as searchable context.
Grouping is policy.
Compare candidates with evidence, not feature matrices
Datadog, Better Stack, Logtail, Axiom, and self-hosted Loki should face one acceptance test. Product pages and plan tables change; the service's correctness questions do not. Export a sanitized fixture containing normal successes, retryable timeouts, permanent rejections, duplicate submissions, and a late callback. Send exactly the same fixture to every candidate, then have two engineers answer the same queries without prior tuning.
| Decision test | Pass condition | Risk |
|---|---|---|
| Trace one notification | Every attempt is ordered and terminal state is explainable | Missing attempts break auditability |
| Group failures | Stable classes form actionable groups | Raw messages fragment incidents |
| Detect regression | A bounded aggregate exposes the rate change | Individual retries flood responders |
| Control sensitive data | Prohibited fields stay outside ingestion and queries | Search expands the compliance boundary |
| Recover exporter failure | Buffer bounds and drops are visible | Silent loss creates false confidence |
| Operate the backend | Upgrades, storage, and recovery have owners | Self-hosting transfers work |
Treat the five names as candidates, not conclusions. Each hosted candidate must be checked against the same ingestion, query, retention, export, and access-control requirements in its current documentation and contract. Loki must be checked against the same outcomes plus the team's ability to operate its storage and query path. The comparison is fair only if engineering labor, on-call ownership, and recovery testing appear beside external charges; a cheap system that cannot answer the audit question has failed before cost comparison begins.
Do not score screenshots. Time the investigation steps, record whether each answer is complete, and retain the query used to obtain it. A fast result that loses duplicate-attempt evidence is worse for this workload than a slower correct result. Conversely, retaining every byte indefinitely is not rigor. Apply field-level classification and documented retention limits.
Correctness beats speed.
Make loss and duplication visible
The application should emit structured events to a local writer or collector without making customer-facing delivery depend on a remote logging endpoint. Buffering needs a finite capacity. When that capacity is exhausted, increment a drop counter and surface it independently, because an invisible logging failure makes every dashboard look healthier.
Workers can emit duplicate records when a process crashes after writing an event but before acknowledging its queue item. Preserve identifiers and deduplicate only in derived views. The raw audit stream shows what happened; the projection prevents a repeated record from inflating logical notification counts. This also makes a correction reproducible: rebuild the projection from source events with a documented ordering rule and deterministic state transition.
Small systems do not need elaborate machinery to enforce the rule. They need an invariant: for one idempotency key, repeated submissions cannot create multiple logical notifications, while every physical attempt receives a distinct attempt ID. Test it under concurrency, retry after timeout, and delayed callback delivery.
Test the boundary.
Roll out the schema before changing the backend
First, emit the new event beside the existing record and validate redaction outside production. Next, build the derived notification-state query and compare it with source-of-truth application state. Then replay the sanitized fixture through each candidate, record query correctness and operational burden, and choose only after acceptance criteria are frozen.
During migration, dual-write for a bounded period and compare counts by low-cardinality outcome, not by eyeballing dashboards. Define cutoff and rollback criteria in advance. After the new path proves it can reconstruct attempts, identify duplicates, expose dropped records, and honor access and retention controls, remove the old path while retaining migration evidence required by audit policy.
The final decision is deliberately unbranded: choose the system that preserves required evidence, produces actionable groups, and has an operating model the team can sustain. Everything else is negotiable.
Top comments (0)