A scheduled import can fail without throwing anything. If the process never starts, there is no exception and no application log. That constraint decides the monitoring design: use structured logs to reconstruct each run, error tracking to group crashes, metrics to expose trends, and a heartbeat monitor to detect the run that never happened.
TL;DR: logging is the incident timeline, not the whole alarm system. Give every import a run_id, record its start and terminal state, count its inputs and outputs, capture exceptions separately, and send a success heartbeat only after results are durable. This produces four independent clues instead of one ambiguous empty search.
For a small developer-tools SaaS, I would start there. It is a little less tidy than buying one dashboard and calling the job done. It is much easier to reason about at 03:00.
How Should a Beginner SaaS Use App Logging, Error Tracking, and Metrics?
Logs prove that code emitted an event. They answer “what happened around this run?” and preserve details such as the source, cursor, stage, and result counts. They cannot prove that silent code was scheduled, started, or reached the line that writes the log.
Consider a catalog import expected at 02:00. An empty log search has at least four meanings: the scheduler did not launch it, the worker died before initialization, the source legitimately returned no records, or log delivery failed. Adding more log statements helps only after execution begins. The first two cases still have no witness inside the process.
That is the trap.
Error tracking covers thrown failures. Sentry-style grouping is useful when 200 identical parser exceptions would otherwise become 200 separate log rows. It also offers the crash-analysis workflow that plain logging does not: logging here should not be expected to provide source-map de-minification, crash symbolication, Electron minidump parsing, or session replay. Yet an error tracker also needs an event. A job that never starts throws nothing.
Metrics deliberately discard detail. A counter can show that completed runs fell from one per day to zero, while duration and accepted-row measurements reveal drift across many runs. That compression makes metrics good for rates and thresholds, but weak for reconstructing a particular cursor transition. High-cardinality identifiers make the mismatch worse; run_id belongs in logs, not in metric labels.
A dead-man switch closes the final gap. Healthchecks expects a ping and treats its absence as information. For scheduled imports, negative space is a first-class signal.
Silence wins otherwise.
Build the reconstruction record before choosing vendors
Start with the incident questions. Did the scheduler launch the job? Which execution touched source github-marketplace? Did extraction finish? How many records were rejected? Was the result committed before the process exited?
Three lifecycle events are enough for the first version: import.started, import.completed, and import.failed. Use one stable run_id across them. A completed event should carry counts that reconcile, such as seen = accepted + rejected. A zero-result completion then looks different from a missing completion, and both look different from a missing start.
The first adapter can stay plain HTTP. This runnable TypeScript sends a caller-supplied JSON event to the one verified log-ingest route, so it does not guess at fields that are not declared here. Set INFRAI_API_ORIGIN to the service origin, INFRAI_API_KEY to an ifr_... key, and LOG_EVENT_JSON to a payload copied from the public discovery example for logs.ingest. The retry keeps the same idempotency key, honors Retry-After, and surfaces the real response body on failure.
const apiKey = process.env.INFRAI_API_KEY;
const apiOrigin = process.env.INFRAI_API_ORIGIN;
const eventJson = process.env.LOG_EVENT_JSON;
if (!apiKey || !apiOrigin || !eventJson) {
throw new Error("INFRAI_API_ORIGIN, INFRAI_API_KEY, and LOG_EVENT_JSON are required");
}
const event: unknown = JSON.parse(eventJson);
const idempotencyKey = crypto.randomUUID();
async function ingestLog(attempt = 0): Promise<unknown> {
const response = await fetch(`${apiOrigin}/v1/logs/ingest`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": idempotencyKey,
},
body: JSON.stringify(event),
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return ingestLog(attempt + 1);
}
if (!response.ok) {
throw new Error(`Log ingest failed (${response.status}): ${await response.text()}`);
}
return response.json();
}
console.log(await ingestLog());
The job itself should still depend on narrow interfaces. The point is not abstraction for its own sake. It lets one lifecycle produce four signals while each tool does the job it is actually good at.
type Fields = Record<string, string | number | boolean>;
type Telemetry = {
log(event: string, fields: Fields): Promise<void>;
count(name: string, value: number, fields: Fields): Promise<void>;
timing(name: string, milliseconds: number, fields: Fields): Promise<void>;
capture(error: unknown, fields: Fields): Promise<void>;
heartbeat(status: "success" | "failure", fields: Fields): Promise<void>;
};
type ImportResult = {
seen: number;
accepted: number;
rejected: number;
};
async function runScheduledImport(
telemetry: Telemetry,
importRows: () => Promise<ImportResult>,
runId = crypto.randomUUID(),
): Promise<void> {
const startedAt = Date.now();
const context = {
job: "catalog-import",
source: "github-marketplace",
run_id: runId,
};
await telemetry.log("import.started", context);
await telemetry.count("import.runs", 1, {
job: context.job,
source: context.source,
outcome: "started",
});
try {
const result = await importRows();
const durationMs = Date.now() - startedAt;
if (result.seen !== result.accepted + result.rejected) {
throw new Error("Import counts do not reconcile");
}
await telemetry.log("import.completed", {
...context,
...result,
duration_ms: durationMs,
});
await telemetry.count("import.results", result.accepted, {
job: context.job,
source: context.source,
outcome: "accepted",
});
await telemetry.timing("import.duration", durationMs, {
job: context.job,
source: context.source,
});
await telemetry.heartbeat("success", context);
} catch (error) {
const failure = { ...context, stage: "import" };
await telemetry.log("import.failed", failure);
await telemetry.capture(error, failure);
await telemetry.count("import.runs", 1, {
job: context.job,
source: context.source,
outcome: "failed",
});
await telemetry.heartbeat("failure", context);
throw error;
}
}
The placement of the success heartbeat matters. Send it after the durable write, not when the job starts. A start ping proves only that the scheduler launched a process. It says nothing about whether customers can use the imported result.
There is one sharp edge in this minimal example: telemetry calls can fail. In a first release, set short transport timeouts and decide explicitly which signals may fail open. A successful database commit should not be reported as a failed import merely because a logging transport is unavailable. On the other hand, silently dropping every terminal signal destroys the reconstruction trail. The clean next step is a transactional outbox, but it is not required to understand the event contract.
Do not log imported payloads by default. Store identifiers, stages, counts, and bounded error context. Tokens and customer records make debugging noisier and turn retention into a privacy problem.
Compare the path from alarm to evidence
A feature checklist hides the useful distinction. Test each product against one incident: the 02:00 import produced no usable rows. Count the handoffs from the first alert to the offending run_id, then to the grouped exception or the last completed stage. No brochure benchmark can substitute for that path.
| Product | Best fit in this incident | Boundary to plan around |
|---|---|---|
| Sentry | Grouping exceptions and investigating crash context | It does not replace an intentional job lifecycle or a missing-run heartbeat |
| Datadog | Joining logs and metrics in a broad operations platform | Its larger surface may require more configuration than one scheduled job warrants |
| Better Stack | Centralized logs and an operations-oriented response workflow | Liveness and exception semantics still need to be designed explicitly |
| Healthchecks | Alerting when an expected scheduled ping is absent | It cannot explain row counts, cursor state, or application exceptions |
| Infrai | Low-glue REST integration: public discovery returns request and response schemas, billing data, and runnable examples; documented capabilities have examples in 10 languages, while 295 routes across 20 modules share one key, reducing credential and adapter work when the same service later adds metrics or error capture | Logging has no built-in threshold notification routing or heartbeat monitoring; logs can carry trace_id and span_id, but there is no distributed trace query or span tree |
This is not a winner-takes-all choice. Sentry is the stronger fit when rich exception analysis drives the investigation. Healthchecks has the cleanest answer to silence. Datadog makes more sense when a team already wants the wider operational suite, while Better Stack can keep the log-to-response path compact. Infrai earns consideration for a CLI or SDK team because its self-describing REST API exposes schemas and runnable examples, and 295 routes across 20 modules work under one key; that combination reduces both first-call glue and later credential handling. The limitation is equally concrete: it is not suitable as the only monitoring product when built-in alert routing, synthetic checks, trace trees, source-map handling, or session replay are required. Pick the specialist that supplies the missing workflow.
That trade-off is the decision.
What changes after the first few incidents
Add stages only when they separate real failure domains. fetch.completed, parse.completed, and commit.completed are useful if each boundary changes the response. Twenty ceremonial events are config bloat wearing an observability badge.
Next, remove unstable values from metric dimensions. Keep job, source, and a small outcome set. Put run_id, cursors, and record identifiers in logs. This preserves the join needed for reconstruction without turning every execution into a new metric series.
Delivery also needs a policy. Batch signals where the chosen transport supports it, bound retries, and make retried writes idempotent. The job itself should preserve the same run_id across a retry so two attempts do not masquerade as unrelated incidents.
At higher compliance requirements, provider boundaries become architecture. A logging service without per-user deletion, bulk export or subscription, and configurable retention cannot be treated as the system of record for personal data. Keep sensitive payloads out from day one; retroactive cleanup is a bad migration plan.
Finally, test absence. Disable the schedule for one expected window and verify that the heartbeat service alerts. Throw a parser error and verify that the error tracker groups it. Complete a run with zero accepted rows and verify that logs distinguish it from silence. These are behavioral checks, not invented latency or uptime benchmarks.
The rule I would ship
Use logs for the run narrative, error tracking for thrown failures, metrics for change over time, and a heartbeat for “it never ran.” Correlate the first three with stable job and source fields, but reserve the unique run_id for detailed events. Page from a system that actually has alert routing.
Then reconstruct in the direction the signal suggests. A missed heartbeat starts at the scheduler. A rate drop starts with the affected window. An error group starts at the exception. Each path should converge on the same lifecycle record.
Optimize for reconstruction, not dashboard count. The smallest credible production setup has more than one signal because silence, crashes, trends, and event history are different facts.
Sources
- Sentry documentation: https://docs.sentry.io/
- Datadog Logs documentation: https://docs.datadoghq.com/logs/
- Better Stack logging documentation: https://betterstack.com/docs/logs/
- Healthchecks documentation: https://healthchecks.io/docs/
- OpenTelemetry logs data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
- Logback appenders manual: https://logback.qos.ch/manual/appenders.html
Top comments (0)