Alert on a missing successful result, not on every exception. Short answer: for a small B2B education product, keep raw server errors as an event stream, group them with a stable failure fingerprint, and join that evidence to a scheduled-import heartbeat. Search, event detail, and resolution still matter, but none can detect a job that never started. The useful four-signal loop is schedule due, run started, result produced, and deadline passed.
This changes the buying question. Rollbar, Bugsnag, Sentry, or a lean internal API can all sit behind the same evaluation harness; the decisive test is whether the complete loop yields one actionable incident when a district roster import goes quiet, without hiding separate affected tenants. I would spend scarce engineering time on that fixture before comparing dashboards or plan tables.
Why doesn't ordinary error grouping catch a silent import?
An exception tracker only receives what the application emits. A scheduler can fail before worker code runs, a queue can retain work beyond the useful window, and a worker can exit without producing a result. In each case, there may be no exception event to group.
Silence is the failure.
The Twelve-Factor logs guidance supplies a clean boundary: the application writes each event to standard output as an unbuffered stream and does not manage log routing itself. That keeps capture separate from interpretation. For this workflow, interpretation means correlating business progress events with errors, rather than asking an error tool to infer that an expected event is absent.
Use four timestamps per scheduled import: dueAt, startedAt, resultAt, and deadlineAt. They are a design model, not an industry standard. Preserve region, tenant, import type, run ID, release, and a sanitized error fingerprint. Do not put student names, email addresses, uploaded row data, or access tokens into the grouping key.
One boundary matters more than it first appears. Region belongs in search and routing metadata, while tenant identity usually should not split the underlying software fault. Otherwise one parser regression affecting 80 schools becomes 80 groups. Yet the incident detail must retain an affected-tenant count, because a single school and a broad outage demand different responses.
Build the smallest useful event loop
The data flow is plain: the scheduler emits import.due; the worker emits import.started; a committed output emits import.result; exceptions emit import.error; and a separate monitor queries overdue runs. That monitor creates or updates one incident per fingerprint and operational scope. Resolution records an operator decision, but a later matching failure must be able to reopen the incident.
Here is a compact TypeScript model with an in-memory implementation. The 15-minute deadline is an explicit product decision for this example, not a universal threshold.
type Region = "us" | "eu";
type ImportEvent = {
kind: "import.due" | "import.started" | "import.result" | "import.error";
runId: string;
tenantId: string;
region: Region;
importType: "roster" | "course";
at: string;
fingerprint?: string;
message?: string;
};
type Incident = {
key: string;
status: "open" | "resolved";
firstSeenAt: string;
lastSeenAt: string;
runIds: string[];
tenantIds: string[];
sample?: ImportEvent;
};
const events: ImportEvent[] = [];
const incidents = new Map<string, Incident>();
function emit(event: ImportEvent): void {
events.push(event);
process.stdout.write(`${JSON.stringify(event)}\n`);
}
function stableErrorKey(event: ImportEvent): string {
const failure = event.fingerprint ?? "missing-result";
return [event.importType, event.region, failure].join(":");
}
function upsertIncident(key: string, event: ImportEvent, now: string): void {
const current = incidents.get(key);
incidents.set(key, {
key,
status: "open",
firstSeenAt: current?.firstSeenAt ?? now,
lastSeenAt: now,
runIds: [...new Set([...(current?.runIds ?? []), event.runId])],
tenantIds: [...new Set([...(current?.tenantIds ?? []), event.tenantId])],
sample: current?.sample ?? event,
});
}
function inspectOverdue(now: Date, deadlineMinutes = 15): void {
const dueEvents = events.filter((event) => event.kind === "import.due");
for (const due of dueEvents) {
const completed = events.some(
(event) => event.runId === due.runId && event.kind === "import.result",
);
const deadline = new Date(due.at).getTime() + deadlineMinutes * 60_000;
if (!completed && now.getTime() > deadline) {
upsertIncident(stableErrorKey(due), due, now.toISOString());
}
}
}
function resolveIncident(key: string, at: string): void {
const incident = incidents.get(key);
if (incident) incidents.set(key, { ...incident, status: "resolved", lastSeenAt: at });
}
This example deliberately leaves transport out. emit can feed a collector, a file-backed forwarder, or a managed ingestion endpoint without changing the event contract. In production, the in-memory collections need durable storage, idempotent writes, retention rules, authorization, and pagination. The monitor also needs an ownership lease so two replicas do not create duplicate notifications.
The fingerprint is intentionally boring. Start with import type, region, and a normalized failure class or code. Never normalize by deleting arbitrary numbers from messages: a course ID may be noise, but an HTTP status or schema version may distinguish two repairs. Prefer an explicit error code created near the failure over message surgery downstream.
How should a small B2B SaaS compare error grouping, search, event detail, and resolve APIs?
Create a fixed replay set and send the same events through every candidate integration, including an internal baseline. This avoids a vague feature contest between Rollbar, Bugsnag, Sentry, and an alternative API. Their interfaces and commercial terms can change; your acceptance cases shouldn't.
| Replay case | Expected incident behavior | Noise failure to reject |
|---|---|---|
| Worker throws three times for one run | One open group with three events | Three notifications |
| Tenants in US and EU hit the same parser fault | Separate regional groups, shared fingerprint visible | One cross-region incident with unclear data routing |
| Scheduler emits due but no start | One missing-result incident after the deadline | No incident because no exception exists |
| Result arrives before the deadline | No alert | A premature missing-result alert |
| Resolved parser fault returns | Reopen or create a clearly linked incident | Append silently to resolved history |
For each candidate, measure the replay output, not a sales checklist. Count groups created, notifications sent, false positives, time from deadline to searchable incident, and fields preserved in event detail. Those are local evaluation measurements. Publish no universal benchmark from them, since ingestion path, sampling, configuration, and region placement all affect the result.
Walk one case all the way through. At 01:00 UTC, the scheduler says a US roster import is due. No start appears. At the chosen 15-minute deadline, the monitor opens a roster:us:missing-result incident and sends one notification. At 01:19, a delayed worker starts; at 01:22, it commits a result. The incident then records the late completion according to the team's policy instead of pretending the deadline miss never happened. Replay the same sequence with three duplicated due events, then with an EU run using the same fingerprint, then with two tenants. The expected group and notification counts must be written down before anyone looks at a candidate's output. This one exercise exposes duplicate handling, regional scope, tenant aggregation, late-event behavior, search latency, and detail quality without inventing a production benchmark.
Search must answer an operator's actual questions: which roster imports missed their deadline in the EU region, which release first carried the fingerprint, and how many tenants are affected? Event detail should reconstruct the progression from due to start to result or error without exposing learner data. Resolution needs an audit timestamp and reason. Fast search with a weak event contract just returns ambiguity sooner.
There is also a cost trap. High-volume retry events can dominate ingestion while adding little diagnostic value. Keep the first useful example, aggregate repetitions, and retain counts. Do not sample away the heartbeat events used to prove completion; losing one can manufacture a false outage.
This design has limits. It isn't suitable when every event payload must be retained verbatim, when operators need full distributed traces rather than import-state evidence, or when the team cannot own a durable monitor and incident state. A hosted tracker can reduce that ownership burden; a self-managed pipeline can offer tighter control but shifts upgrades, storage, and on-call work back to the team. The trade-off is operational responsibility versus control, not a universal winner. Rollbar, Bugsnag, and Sentry should each be judged against the same replay and data-boundary requirements rather than assumed equivalent.
No shortcut fixes a weak contract.
Treat flags and releases as evidence, not causes
Import behavior often changes behind a feature flag: a new CSV parser, a revised mapping rule, or a tenant-specific rollout. Record the flag key and evaluated variant on the run event when that metadata is available. An experimentation system such as GrowthBook is one possible source of that context, but the incident model should accept generic flag metadata so diagnosis is not coupled to one control plane.
Correlation is not causation. A new variant appearing beside failures narrows the investigation; it does not prove the flag caused them. The replay test should cover both variants, and rollback policy should be defined before rollout. This is particularly valuable for a solo builder: the incident detail can answer “what changed?” without requiring a second person to remember deployment history.
Keep it lean.
Operate the loop without creating another noisy system
Start deployment in shadow mode. Generate incidents and inspect them, but suppress paging for at least one complete import cycle chosen for the application's schedule. Compare expected runs with produced results, then tune deadlines from observed completion distributions and business commitments. A hard-coded 15 minutes is only a runnable starting point.
Before enabling notifications, verify that every scheduled run has one stable ID across scheduler and worker boundaries, duplicate events are harmless, clocks are synchronized closely enough for the deadline, and delayed results close or annotate missing-result incidents predictably. Confirm that US and EU routing follows the organization's data-handling requirements. Then test access controls with an account that should see metadata but not payload contents.
Day to day, review unresolved groups by affected tenants and age, inspect the sample event plus adjacent progress events, and record a resolution reason. Weekly, look for fingerprints that repeatedly reopen, deadlines that create false alarms, and retry storms that inflate storage. After a schema or scheduler change, replay the fixture before shipping. This is mundane work, which is exactly why it survives vendor changes.
The final selection should be the implementation that passes this loop with the least operational burden you can sustain: reliable missing-result detection, predictable grouping, region-aware search, sufficient detail, and auditable resolution. Choose on signal quality under your own replay, not on the length of a feature matrix.
Top comments (0)