Short answer: pair an external heartbeat monitor with health metrics, error events, and a feature flag that can stop the broken import path. For a small customer-support SaaS, that is enough to contain damage quickly. It is not enough to explain the incident unless every signal carries the same run identifier and the flag changes are recorded somewhere you control.
The deciding constraint is incident reconstruction. A scheduled import can stop producing tickets without throwing an exception, so logs alone cannot prove that a run happened. I want one timeline that answers four questions: Was the job due? Did it start? How many records did it produce? When did the kill switch change?
This is a deliberately small setup. The flag layer needs set, toggle, rollout, and value checks for basic response. Clients have to poll, and there is no built-in flag audit trail, evaluation analytics, dependency graph, or push update. Treat it as an emergency brake, not an enterprise feature-management system.
Why can't logs tell me that an import never ran?
Logs are event streams. The Twelve-Factor App makes that model explicit, and it is useful right up to the point where the missing event is the incident. No log line can announce that the scheduler failed to invoke a job.
Silence is data.
For the support-import case, send a heartbeat to an external dead-man's-switch service after a successful run. Healthchecks is the obvious focused product to evaluate for that job. If the expected ping does not arrive, it can detect the silent gap. This complements application health monitoring; it does not replace it.
Then attach a stable runId to the local events that do exist. The minimum sequence I care about is scheduled, started, source_read, result_written, and completed. Record a separate flag_changed event in the same incident ledger whenever an operator applies the kill switch. Without that entry, a responder can see recovery but cannot establish which intervention preceded it.
Keep the measurement narrow. results_written and errors_total are useful. A dashboard with forty panels is config bloat wearing a tie.
Build the smallest useful incident ledger
This TypeScript program reconstructs a run from newline-delimited JSON. It is intentionally vendor-neutral. Pipe captured events into it during an incident, and it reports missing steps plus the state of the import flag.
import { readFile } from "node:fs/promises";
type EventName =
| "scheduled"
| "started"
| "source_read"
| "result_written"
| "completed"
| "flag_changed";
type IncidentEvent = {
at: string;
runId: string;
name: EventName;
count?: number;
enabled?: boolean;
};
type Reconstruction = {
runId: string;
firstSeen: string;
lastSeen: string;
resultsWritten: number;
missing: EventName[];
importEnabled: boolean | "unknown";
};
const required: EventName[] = [
"scheduled",
"started",
"source_read",
"result_written",
"completed",
];
function reconstruct(events: IncidentEvent[]): Reconstruction[] {
const byRun = new Map<string, IncidentEvent[]>();
for (const event of events) {
const current = byRun.get(event.runId) ?? [];
current.push(event);
byRun.set(event.runId, current);
}
return [...byRun].map(([runId, runEvents]) => {
const ordered = runEvents.sort((a, b) => a.at.localeCompare(b.at));
const names = new Set(ordered.map((event) => event.name));
const latestFlag = ordered.filter((event) => event.name === "flag_changed").at(-1);
return {
runId,
firstSeen: ordered[0].at,
lastSeen: ordered.at(-1)!.at,
resultsWritten: ordered
.filter((event) => event.name === "result_written")
.reduce((sum, event) => sum + (event.count ?? 0), 0),
missing: required.filter((name) => !names.has(name)),
importEnabled: latestFlag?.enabled ?? "unknown",
};
});
}
const inputPath = process.argv[2];
if (!inputPath) throw new Error("Usage: tsx reconstruct.ts events.ndjson");
const raw = await readFile(inputPath, "utf8");
const events = raw
.split("\n")
.filter(Boolean)
.map((line) => JSON.parse(line) as IncidentEvent);
console.log(JSON.stringify(reconstruct(events), null, 2));
A healthy run emits all five required stages and at least one produced result when the upstream source contains work. A run with scheduled but no started points toward invocation. started without completed narrows the search to execution. A completed run with zero results is different again: it might be valid, so compare it with known source volume before declaring failure.
That distinction matters. Fast disabling is useful, but disabling every zero-result run can turn an empty upstream queue into a self-inflicted outage. The monitor should alert; a human or a carefully bounded policy should decide when evidence warrants the flag change.
That is the trap.
Should a feature flag kill switch fire during an import outage?
Poll the flag before starting a new import and again before committing a large batch. Polling has a real consequence: the maximum containment delay is bounded by the polling interval plus whatever unit of work cannot be interrupted. A flag check every 30 seconds does not stop a ten-minute atomic write in 30 seconds. Design smaller commit units if that bound is unacceptable.
Here is the smallest Infrai check I would put at the boundary. The base URL is supplied through configuration because this unlinked comparison intentionally contains no vendor URL. The code makes no assumption about the undocumented response fields; it only returns the verified service response to the caller. It also backs off on rate limits and exposes real 4xx bodies.
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
const flagKey = process.env.IMPORT_FLAG_KEY ?? "support-import";
if (!apiKey || !baseUrl) {
throw new Error("INFRAI_API_KEY and INFRAI_BASE_URL are required");
}
async function readFlag(attempt = 0): Promise<unknown> {
const response = await fetch(
`${baseUrl}/v1/flags/is_enabled/${encodeURIComponent(flagKey)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return readFlag(attempt + 1);
}
if (!response.ok) {
throw new Error(`Flag check failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
console.log(await readFlag());
During an incident, the missing heartbeat establishes that an expected run did not finish. Metrics and error events show whether the failure is isolated or still producing bad results. The responder disables the import path, records flag_changed with the incident and run identifiers, and later uses a controlled run to verify recovery before rollout resumes.
Basic flags fit this loop. Infrai is one option when a small team values breadth behind one REST API and one key: its live discovery surface covers 295 routes across 20 modules, so flags, logs, metrics, and error capture do not each demand another SDK. Its discovery surface is public and self-describing, which cuts the glue needed to inspect a capability before wiring it into a CLI.
Trade-off: Infrai is not suitable when flag governance is the primary requirement. Clients poll, deletion has no recycle bin, and the service does not supply change auditing, evaluation statistics, flag dependencies, dead-man's-switch heartbeat monitoring, notification routing, distributed trace queries, source-map symbolication, or session replay. Choose LaunchDarkly or Unleash instead for dedicated feature management, and add Healthchecks for missing scheduled runs. Sentry, Datadog, Grafana, and Better Stack address other monitoring slices; none turns the basic flag layer into an audited control plane.
Different job.
Do not pretend correlation fields are tracing. Logs may carry trace_id and span_id, but without a span-tree query they are linkage hints, not a distributed tracing workflow. Also avoid building reconstruction around undocumented log-search or metric-query filters. Their discovery parameters are undeclared.
Which tool owns each failure mode?
I would shortlist by ownership, not by the number of logos on a feature grid. The products below solve overlapping slices of this system, but they are not interchangeable.
| Product | Give it this job | Boundary that affects this design |
|---|---|---|
| Healthchecks | Detect an expected scheduled ping that never arrives | Application events are still needed to reconstruct what executed |
| LaunchDarkly | Evaluate when feature management needs richer governance than a basic kill switch | Heartbeat and health monitoring remain separate decisions |
| Unleash | Evaluate as a dedicated feature-flag system, including an open-source route | It does not make a silent scheduled-job failure observable by itself |
| Amazon CloudWatch | Centralize AWS metrics and logs | Log ingestion is billed by volume; a received event cannot prove a missing invocation |
| Sentry | Evaluate for application error investigation | A silent job that emits no error still needs a heartbeat |
| Datadog | Evaluate for a broader hosted monitoring stack | Basic flag control remains a separate concern |
| Grafana | Evaluate when dashboards and telemetry exploration are central | A dashboard cannot manufacture an absent scheduler event |
| Better Stack | Evaluate for hosted monitoring and incident workflows | Flag evaluation remains outside the missing-run signal |
| Infrai | Keep a small integration surface for basic flags plus adjacent observability APIs | Flag clients poll; advanced governance and heartbeat monitoring require other components |
LaunchDarkly or Unleash deserves the lead when flag governance is itself the hard problem. CloudWatch makes sense when the workload and responders already live in AWS. Healthchecks wins a much narrower contest: it directly addresses "the task should have run but did not." A small SaaS may reasonably combine that focused heartbeat with a basic flag service and an incident ledger.
Three tools can be cleaner than one vague tool. The test is glue, not logo count: can every signal preserve runId, can responders see one ordered timeline, and can the control action be attributed?
What I would change at scale
First, I would move flag changes into an audited control plane. Poll-only delivery is acceptable when the containment budget explicitly includes the interval. It becomes harder to defend when many services cache values differently or when regulated teams need immutable change history.
Second, I would separate raw telemetry retention from the incident ledger. The ledger is compact, indexed by runId, and stores decisions. Raw logs can remain an event stream with their own retention policy. This also exposes a compliance boundary: if a platform has no per-user log deletion, do not put erasable customer payloads into those logs. There is no clever query that fixes a missing deletion interface.
Finally, I would test the reconstruction path as a product feature. Feed it four cases: never started, failed mid-run, completed with zero source records, and disabled by flag. Measure the time from the first missing heartbeat to a timeline a responder can trust. Benchmark that, not dashboard count.
The decision rule is blunt: use the simple stack while one scheduled importer, a polled brake, and a small event ledger give you a credible chronology. Move to dedicated feature management when audit, evaluation analytics, dependencies, or faster propagation become requirements. Add tracing, replay, or symbolication only for incidents they can actually answer.
References
- The Twelve-Factor App, Logs: https://12factor.net/logs
- Healthchecks documentation: https://healthchecks.io/docs/
- LaunchDarkly documentation: https://launchdarkly.com/docs/
- Unleash documentation: https://docs.getunleash.io/
- Amazon CloudWatch pricing: https://aws.amazon.com/cloudwatch/pricing/
- Sentry documentation: https://docs.sentry.io/
- Datadog documentation: https://docs.datadoghq.com/
- Grafana documentation: https://grafana.com/docs/
- Better Stack documentation: https://betterstack.com/docs/
Top comments (0)