The cheapest log search system is the one that preserves enough evidence to explain why a marketplace notification failed without making every retry look like a new incident. Fix the event shape, cardinality, and retention tiers first; then compare self-hosted search, a managed search cluster, and a hosted logs API against the same workload. Otherwise the vendor comparison measures noise.
TL;DR: emit one structured delivery-attempt event per provider call, keep stable identifiers searchable, collapse retries when calculating alerts, and separate short-lived diagnostic payloads from longer-lived outcome records. For a small team, the decisive trade-off is signal quality versus operational noise. Hosting model comes after that.
What question must the logs answer?
A marketplace notification service has a deceptively simple job: accept an intent, ask a delivery provider to send it, retry failures, and record the final outcome. Search gets messy because one buyer action can generate several attempts. An alert on every failed attempt pages on recoverable retries; an alert only on terminal failure can hide provider degradation until the retry window closes.
Define the investigation question before selecting storage: "Which delivery intents failed permanently, for which channel and failure class, during this release window?" That wording forces a useful distinction between an intent and an attempt. It also keeps message contents out of the primary index. GDPR Article 5 requires personal data to be adequate, relevant, and limited to what is necessary, so indexing an email address or full message body by default needs a purpose stronger than convenience.
The useful flow is small enough to explain plainly. The service creates a stable notificationId, records every provider attempt with an attemptNumber, normalizes the provider response into a bounded failure class, and emits a terminal outcome when retry policy finishes. Searchable storage receives the normalized envelope. A separately controlled diagnostic tier may receive a redacted provider response for a shorter period. Alerts operate on terminal outcomes and rates, while an investigator can still expand one intent into its attempts.
One intent. Several attempts. One final outcome.
Build the evidence before buying the index
This runnable TypeScript example models that boundary. It deliberately avoids a logging SDK. The output can go to standard output, an agent, a queue, or an HTTP collector without changing the event contract.
type Channel = "email" | "sms" | "push";
type FailureClass =
| "authentication"
| "rate_limit"
| "invalid_destination"
| "provider_unavailable"
| "unknown";
type AttemptEvent = {
event: "notification.delivery_attempt";
occurredAt: string;
notificationId: string;
marketplaceId: string;
channel: Channel;
attemptNumber: number;
release: string;
outcome: "delivered" | "failed";
failureClass?: FailureClass;
providerStatus?: number;
durationMs: number;
};
type ProviderResult = { ok: boolean; status: number; code?: string };
function classifyFailure(result: ProviderResult): FailureClass {
if (result.status === 401 || result.status === 403) return "authentication";
if (result.status === 429) return "rate_limit";
if (result.code === "INVALID_DESTINATION") return "invalid_destination";
if (result.status >= 500) return "provider_unavailable";
return "unknown";
}
function makeAttempt(
result: ProviderResult,
attemptNumber: number,
durationMs: number
): AttemptEvent {
const base = {
event: "notification.delivery_attempt" as const,
occurredAt: new Date().toISOString(),
notificationId: "ntf_7f31",
marketplaceId: "mkt_204",
channel: "email" as const,
attemptNumber,
release: "2026-10-04.1",
durationMs
};
return result.ok
? { ...base, outcome: "delivered" }
: {
...base,
outcome: "failed",
failureClass: classifyFailure(result),
providerStatus: result.status
};
}
const attempts = [
makeAttempt({ ok: false, status: 429 }, 1, 184),
makeAttempt({ ok: false, status: 429 }, 2, 391),
makeAttempt({ ok: true, status: 202 }, 3, 146)
];
for (const attempt of attempts) {
process.stdout.write(`${JSON.stringify(attempt)}\n`);
}
The three records describe one notification, not three customer-impacting failures. That is the trap. This example has exactly three attempts: two responses with status 429, followed by one accepted response with status 202. Counting lines where outcome=failed produces a failure count of two; grouping by notificationId shows eventual delivery. Keep both views because they answer different questions: attempt failures expose provider pressure, while terminal outcomes describe user impact. The trade-off is explicit. Alerting on attempts is early but noisy; alerting on terminal outcomes is quieter but later. I'd store both and page from neither in isolation. A rate of recovered retries belongs beside the terminal-failure rate, so a provider slowdown is visible without turning each recovered notification into customer impact.
Field design controls cost and usefulness more directly than a brand name does. Keep channel, failureClass, release, and perhaps a coarse tenant identifier indexed because they bound common investigations. Keep request IDs and notification IDs searchable if point lookups are routine. Do not turn raw error messages, destination addresses, or message bodies into labels. They are high-cardinality, may contain personal data, and rarely improve an aggregate alert.
One detail deserves restraint: durationMs is a measurement, while buckets are a query choice. Preserve the integer in the event and derive buckets downstream. Baking labels such as fast and slow into emission code makes a later threshold change look like a schema migration.
Should a small app use self-hosted or hosted log search?
Do not compare a self-hosted stack, a managed search cluster, and a hosted logs API by their smallest advertised bill. Send the same representative events through each candidate, then run the same investigations. Public cloud pricing pages can expose ingestion as a metered dimension; the Amazon CloudWatch pricing page is one concrete example. Event volume and retained bytes are therefore design inputs even when the final platform differs.
| Operating model | You operate | Main noise risk | Evidence to test |
|---|---|---|---|
| Self-hosted search | upgrades, storage, backups, query capacity | infrastructure alerts can bury app alerts | restore drill, ingestion backlog, query behavior during retry bursts |
| Managed search cluster | schema, retention, capacity policy, access | permissive indexing expands fields and spend | mapping growth, rejected writes, retention deletion |
| Hosted logs API | event contract, quotas, export path, access | convenience encourages unfiltered payloads | throttling behavior, export completeness, query semantics |
These are responsibility boundaries, not rankings. A solo operator may accept less infrastructure control to reduce maintenance, or accept more operations to keep data placement and query behavior under direct control. The answer changes with on-call capacity, compliance boundaries, and the amount of diagnostic detail the service genuinely needs. It shouldn't change because a demo dataset contained only successful deliveries.
Use a replayable evaluation set: successful first attempts, a delivery that succeeds on attempt three, an invalid destination, authentication rejection, rate limiting, and an unknown response. Add a release marker and two marketplace identifiers. Then ask each system for terminal failures by channel, all attempts for one notification, failure-class rates around a release, and records eligible for deletion. Record whether the query is accurate before recording how quickly it returns.
Fast nonsense is still nonsense.
Retention is part of the schema
A single retention period is easy to configure and hard to justify. Outcome envelopes are compact and useful for longer trend windows. Detailed provider responses are larger, less consistently structured, and more likely to carry data that shouldn't sit in a broad search index. Separate them at ingestion instead of hoping a future cleanup query can identify every sensitive field. A practical policy can use three clocks without pretending the numbers are universal: keep attempt envelopes for the investigation window the support process actually uses, keep terminal outcomes long enough to compare releases and delivery health, and keep redacted diagnostic payloads for the shortest approved debugging window. The exact durations belong in a written policy tied to operational and legal requirements, not copied from a vendor default. Deletion must be testable as well. Select a known notification ID, verify which tiers contain it, execute the normal expiry or deletion process in a staging environment, and verify that all expected copies disappear. Inspect archives and exports too. Data minimization fails if the primary index expires on schedule while an unrestricted export lives indefinitely.
This is where a cheap-looking system can become expensive in attention: one store may reduce configuration but force broad access and blunt retention; several tiers improve control but create more paths to test. Choose the fewest tiers that express the required access and retention boundaries. No fewer.
Three clocks, not three arbitrary copies.
Operate the decision, not the demo
Before deployment, freeze the event contract with schema tests and fixtures for all six evaluation cases. Make emission failure non-blocking for notification delivery, but expose dropped or rejected telemetry through a separate health signal. Roll out the new envelope beside the old one for a bounded migration window, compare terminal counts, and remove the duplicate path on a named date. An open-ended dual-write doubles noise and makes every discrepancy ambiguous.
During operation, review three things together: terminal delivery failures, failed attempts that later recovered, and telemetry loss. A rise in recovered retries can warn about provider pressure without overstating customer impact. Telemetry loss is different; it lowers confidence in both measures and should be visible as such.
The operational checklist is prose because the dependencies matter. First, confirm that every intent has one stable identifier and every call increments its attempt number. Next, verify that failure classes are bounded and unknown values remain inspectable through controlled diagnostics. Confirm access rules and expiry independently for each tier. Re-run the saved investigation queries after schema or retention changes. Finally, test export and deletion before treating portability or compliance as solved.
Choose the operating model only after this evidence pipeline works. The useful comparison is concrete: which option returns correct incident evidence, stays understandable during retry bursts, meets the data boundary, and consumes an acceptable amount of operator time? That decision survives changing price sheets. A ranking based on an empty index does not.
Sources
- GDPR Article 5, principles relating to processing of personal data: https://gdpr-info.eu/art-5-gdpr/
- Amazon CloudWatch pricing: https://aws.amazon.com/cloudwatch/pricing/
Top comments (0)