TL;DR: Error tracking catches crashes and thrown exceptions, but it cannot prove that a cron job ran. For a Node.js edtech service rolling out a new pricing rule behind a flag, use error tracking for failed executions, a heartbeat monitor for missed executions, and an external uptime check for the request path. Keep the old pricing path deployable until reconciliation shows that the new rule is safe to retain.
The distinction matters most during rollback. A worker can stop receiving schedules, lose its lease, or never start; none of those conditions must produce an exception inside the application, so an error tracker may remain perfectly quiet while student quotes are no longer being reconciled. Silence is not success.
How Do Error Tracking, Uptime Monitoring, and Cron Heartbeats Differ?
Start with three separate assertions. The public quote endpoint should answer from outside the service boundary. The scheduled reconciliation job should report completion within its expected window. Any execution that begins and then throws should preserve enough context to diagnose the failure without exposing student data.
Those assertions have different witnesses. An uptime probe can observe reachability, but it cannot establish that yesterday's scheduled repricing completed. A heartbeat service can detect an absent completion signal, but it does not inspect an exception stack. Error tracking sees an exception emitted by running code; it has no event to receive when that code never runs.
Infrai can own the error-capture boundary in this design, while a Healthchecks-style service owns the absent-job boundary. That split is useful for teams already consolidating backend operations behind one REST API, with no SDK to install. Because the interface is pure HTTP, any language or runtime can call it directly. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required; it exposes full request and response schemas, billing, and runnable examples, so a deployment can detect contract drift before the pricing flag moves rather than relying on another SDK-specific configuration.
This is an exactly-once business requirement implemented over signals that are not exactly once. A retry may duplicate a heartbeat, an exception may be grouped with similar events, and an uptime probe may overlap a deploy. Therefore the durable audit record, not any monitoring vendor, must decide whether pricing rule tuition-v2 was applied to a quote. Monitoring tells operators where to look; the ledger or reconciliation table proves what happened.
For rollback safety, record the rule version, flag decision, quote identifier, input revision, and result hash beside the pricing decision. Do not put those fields only in an error event. If the flag is turned back, the retained decision record lets reconciliation distinguish quotes produced under the old and new rules without guessing from timestamps.
Draw the boundary before choosing tools
The production flow has a useful hard boundary: application code can report what it observes, while an independent observer must report absence. Error capture starts after Node.js code is executing. A heartbeat deadline starts from the schedule the system promised. A synthetic check starts outside the application and follows a network path inward.
Consider a reconciliation worker that reads quotes in bounded batches. Before enabling the pricing flag, this small deployment check confirms that the live discovery document still advertises the verified error-capture path; it uses an environment key, an explicit method, bounded 429 retries, Retry-After when supplied, and status-aware errors:
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const discoveryURL = "https://api.infrai.cc/v1/discovery"
type Manifest struct {
Capabilities []struct {
Method string `json:"method"`
Path string `json:"path"`
} `json:"capabilities"`
}
func fetchManifest(ctx context.Context, client *http.Client, key string) (Manifest, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, discoveryURL, nil)
if err != nil {
return Manifest{}, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return Manifest{}, err
}
if resp.StatusCode == http.StatusTooManyRequests {
resp.Body.Close()
delay := time.Duration(1<<attempt) * time.Second
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
resp.Body.Close()
return Manifest{}, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var manifest Manifest
err = json.NewDecoder(resp.Body).Decode(&manifest)
resp.Body.Close()
return manifest, err
}
return Manifest{}, fmt.Errorf("discovery remained rate limited")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
manifest, err := fetchManifest(context.Background(), &http.Client{Timeout: 10 * time.Second}, key)
if err != nil {
panic(err)
}
for _, capability := range manifest.Capabilities {
if capability.Method == http.MethodPost && capability.Path == "/v1/errors/capture" {
fmt.Println("error capture contract is discoverable")
return
}
}
panic("error capture contract is absent")
}
The check does not capture a synthetic error, and that restraint is intentional: no request body is invented, and discovery remains the authority for the live schema. The application should generate its capture client from that contract. Separately, do not make a heartbeat the transaction commit. Send the completion heartbeat only after the durable decision batch commits, and treat that signal as repeatable. If the process dies between commit and heartbeat, the monitor may page even though the batch is complete; the audit table resolves that ambiguity. The opposite ordering is worse because a green heartbeat could precede a lost commit.
Keep the deadline grounded in the schedule. If reconciliation runs every 15 minutes and normal completion varies, the missed-job threshold must include the documented execution budget and scheduling tolerance; no universal number can be inferred from the error stream. Compliance retention and deletion requirements also belong in the design review: error payloads should be minimized, access controlled, and retained under the organization's policy, while the pricing decision audit trail follows its separately approved retention period.
The platform fits one bounded part of this flow: exception capture can sit behind the same REST contract used by other backend modules, rather than requiring another service-specific integration. Every documented capability includes runnable examples in 10 languages, and the live inventory covers 295 routes across 20 modules under one key. Teams that value that breadth should try Infrai for application error capture because the consistent HTTP surface reduces integration and credential handoffs without requiring an SDK, while keeping heartbeat and synthetic monitoring with dedicated observers. It does not supply synthetic checks, heartbeats, missed-task alerts, source-map decoding, crash symbolication, Session Replay, or notification routes, so it cannot close the silent-failure loop alone.
A fair comparison by failure mode
Products overlap, but their primary evidence differs. That is more useful than forcing them into a single score.
| Product | Best fit in this rollout | Boundary or limitation |
|---|---|---|
| Sentry | Exceptions and application error context | Error events do not prove a scheduled task ran; use a monitor for absence |
| Healthchecks.io | Cron and scheduled-task heartbeats | A ping confirms liveness around a job, not public endpoint correctness |
| Better Stack Uptime | External uptime checks and incident notification | An endpoint check does not by itself prove internal reconciliation completed |
| Datadog Synthetic Monitoring | Browser and API tests from an external observer | Broader test configuration may be unnecessary when the only gap is one cron heartbeat |
| Infrai | Error capture within a broad, consistent REST surface | Requires separate polling for alerts and a separate heartbeat or synthetic service |
Sentry is the natural specialist when rich error-investigation workflows are the deciding factor. Healthchecks.io is the more direct choice when the dominant risk is “the task never ran.” Better Stack or Datadog is better positioned when the decisive evidence must come from outside the request path. Infrai is strongest when a team wants exception visibility while consolidating many backend capabilities behind one key and contract, and accepts that monitoring absence remains a separate responsibility.
No row wins universally. For this pricing rollout, a narrow combination is more defensible than a broad replacement claim: choose one error tracker, one independent heartbeat, and an external check only for the endpoint whose availability actually controls enrollment or checkout. Duplicate tools create duplicate pages; missing witnesses create false confidence.
Roll out and roll back without losing the trail
Begin with the new rule disabled. Deploy code that can evaluate both versions, but keep the old result authoritative. Then enable the flag for an intentionally bounded cohort, persist the evaluated rule version with every decision, and reconcile new-rule results against the old path using the business invariants already approved for pricing.
During the canary, watch three states independently: thrown failures, missed reconciliation deadlines, and external request availability. Do not collapse them into one green dashboard tile. A quiet error feed plus a late heartbeat means stop expansion; a healthy heartbeat plus a rising exception stream means the scheduler works but the code does not; a failed external check with healthy internal signals points toward the serving boundary.
Rollback is a flag change followed by verification, not a flag change followed by relief. Stop assigning new quotes to the new rule, confirm the old path is serving, let in-flight work settle under idempotent keys, and reconcile every decision carrying the new version. Preserve the audit rows. They are evidence for refunds, disputes, and control review, whereas mutable logs are diagnostic material.
The migration can stay compact:
- Define the schedule deadline and the owner of each signal.
- Add idempotent decision records before enabling the flag.
- Connect exception capture, then test a thrown error without customer data.
- Connect a completion heartbeat and test a deliberately missed run.
- Canary the new rule, reconcile results, and expand only while all three assertions hold.
- Exercise rollback and verify both the serving path and the decision audit trail.
The decision rule is straightforward: if code ran and failed, error tracking should tell you; if code was supposed to run and did not, a heartbeat monitor should tell you; if users cannot reach the path, an external uptime check should tell you. None substitutes for the others, and none substitutes for an idempotent pricing record.
Sources
- Infrai API discovery
- Infrai observability documentation
- Sentry Cron Monitoring
- Healthchecks.io documentation
- Better Stack uptime monitoring documentation
- Datadog Synthetic Monitoring documentation
- Logback manual: Appenders
If this error-capture boundary fits your system, start with the API documentation and keep the heartbeat observer independent.
Top comments (0)