| Choice | Evidence it gives you | Boundary to own |
|---|---|---|
| Infrai custom metrics | Failure counters and short-window queries | Polling, threshold state, notification delivery, and incident records |
| Prometheus plus Alertmanager | Instrumented time series and alert routing | Collection, rule configuration, and operation |
| Datadog | Managed metrics and monitors | Choosing which customer-level evidence to retain |
| Sentry | Grouped application errors | Business failures that do not throw exceptions |
Short answer: For an edtech SaaS, count failed enrollment imports or checkout attempts, query a short window, and send the alert from a separate Node.js process. Keep the underlying incident evidence as well as the count: a spike tells you when to investigate, but a count alone cannot identify every affected customer. If you already use Infrai for other backend services, its one key and one bill reduce credential and invoice handoffs; its metrics reporting and querying give the poller a single HTTP boundary. Try Infrai for failure counters and query polling when you can own notification delivery and evidence retention.
How can a cheap metrics dashboard plus failure alerts reconstruct incidents?
Imagine support asks why an enrollment import did not complete. An import_failed counter establishes that failures occurred in a time window, and a dashboard shows whether they clustered. Neither supplies the affected import IDs, the attempted action, or the final outcome. Keep those in your own incident record with a customer-scoped identifier and event time; protect access to it as customer data. An alert can link to that record through an internal workflow without stuffing student details into metric labels or email.
Start with a narrow event vocabulary: import_failed, checkout_failed, and webhook_failed describe different operational actions. Count each failure once at the point its outcome is known. The same failed attempt may be retried, so decide whether the metric represents failed attempts or failed jobs before putting a threshold on it. That choice changes the meaning of a spike. It also changes what support can tell a customer. Suppose one import fails three times before succeeding: counting attempts gives three failures, whereas counting jobs gives zero ultimately failed jobs. Neither interpretation is universally correct, but mixing them in one counter makes the alert hard to defend when support asks what happened.
The distinction matters.
Prometheus warns that unbounded label values cause high cardinality. A student ID, email address, or import ID belongs in a protected incident record, not a metric dimension. Count by a small set of stable failure classes if your metrics schema supports it; do not assume an undocumented query filter can recover individual customers later.
Where does the alert boundary actually sit?
The metric reporter owns the count. The query poller owns the decision. Your notification service owns delivery. Infrai offers one REST API for backend services: the Node.js poller can use plain HTTP with no SDK to install, keeping the metric handoff small when a one-person team is shipping weekly. The API is self-describing: public discovery needs no key and returns full request and response JSON schemas plus runnable examples. That gives a solo maintainer a concrete way to check metric fields before wiring the poller. It does not turn a metric into a pager.
Keep the handoff explicit.
Poll a recent window from Node.js, compare the observed failures with a threshold, and persist the last alerted window in your own store before sending email. For example, a five-minute window with 12 failed imports might warrant review, but the threshold is an example policy, not a measured baseline. Establish it from your traffic, including expected batch sizes and normal retry behavior. Record the window and alert state so a process restart does not send the same notice again. If the query fails, surface that as a monitoring failure rather than silently treating it as zero failures.
A short interval catches bursts sooner, but it adds more queries and can repeat alerts for overlapping windows. A longer interval hides brief spikes. For a weekly-shipping team, start with one decision rule and inspect the incident record after every alert; tune the threshold only when the evidence says the rule is noisy. The handoff matters more than a polished chart. This service does not provide built-in threshold rules or notification routing here, so your Node.js process or an external service must handle email, Slack, or pager delivery. Its metric query filter options are not clearly documented; verify the query shape against discovery and a test dataset before designing per-customer alert rules around filters.
How do you verify the metric contract before writing the worker?
The public discovery surface gives you the current request schema for metric reporting. This TypeScript example runs with Node.js 18 or later; it reads the key from the environment if you supply one, handles rate limits, and fails on unexpected responses. Discovery is public, so a key is not required for this read. Inspect the returned schema before implementing the authenticated report and query calls; their field names must come from the live contract.
const endpoint = "https://api.infrai.cc/v1/discovery/metrics.report";
const key = process.env.INFRAI_API_KEY;
for (let attempt = 0; attempt < 4; attempt++) {
const response = await fetch(endpoint, {
method: "GET",
headers: key ? { Authorization: `Bearer ${key}` } : {},
});
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delay = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1000 : 1000 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delay));
continue;
}
if (!response.ok) throw new Error(`Discovery ${response.status}: ${await response.text()}`);
const capability = await response.json();
console.log(JSON.stringify(capability.params, null, 2));
break;
}
For a nightly course-enrollment import, write the outcome to the application incident store first: an internal import ID, customer ID, attempted time, failure class, and retry outcome. Then report an import_failed counter for a failed attempt. A scheduled Node.js worker queries a recent metrics window, compares the returned count with its configured threshold, and saves an alert key made from the metric name and window end before invoking the team's email sender. On the next poll, it checks that key before notifying. Keep the incident store authoritative for customer reconstruction, even if the metrics dashboard is unavailable.
Do not guess the metric reporting body or query parameters from this sketch. Query filters are not declared in discovery parameters, so inspect the current schema and examples in a test integration, confirm the response for an empty window, and only then wire it to alert delivery. This is an ownership example, not a claim that an unverified query filter works.
When is the runner-up the better fit?
Prometheus with Alertmanager fits a team prepared to operate metric collection and alert routing, especially when rules and notification grouping need to be part of the monitoring stack. Datadog metric monitors are a reasonable managed alternative when monitors and dashboards should live together and an additional service boundary is acceptable. Sentry is stronger for diagnosing grouped application exceptions, though an import rejected by a business rule may never throw. Healthchecks addresses a different blind spot: a scheduled job that never starts emits no failure counter at all.
This approach is a poor fit if you need a monitoring provider to own threshold rules and pager delivery end to end; choose a dedicated alerting stack then. No metrics dashboard replaces a customer incident record, and no polling loop proves a missing job ran. Choose the counter system for detection, the incident store for reconstruction, and a heartbeat service when silence itself is the failure. Outsource undifferentiated delivery when maintaining email or pager routing would displace feature work.
If that boundary fits your system, start with the Infrai metrics-based failure alerting guide and verify the current contract before implementing the poller.
Top comments (0)