For a small SaaS error monitoring setup, choose a basic errors API when fast capture, grouped search, and in-app resolution are enough. Choose Sentry or Rollbar when notification routing and deeper production debugging must arrive as part of the product rather than as engineering work around it.
TL;DR: For a small B2B SaaS comparing an AI experiment across tenant cohorts, I would make tenant_cohort, experiment_variant, and model cost part of the application's own error context first. Then I would select the monitor. Cost attribution stays portable, while the monitoring decision comes down to a sharper question: do we merely need searchable grouped failures, or do we need alerts, source-map processing, distributed traces, and session replay too?
That ordering matters. A polished error dashboard cannot answer whether variant B failed more often per dollar of model usage if the application never preserved the cohort and cost dimensions. Instrument the decision, then buy the debugging depth.
What error monitoring setup should a small Node.js SaaS use?
Picture one request from a tenant in the mid_market cohort. It enters an experiment, calls a model, records the returned per-call cost metadata, and either succeeds or raises an exception. Before sending the exception anywhere, the application creates a small, vendor-neutral record containing the cohort, variant, operation, cost, exception type, and a stable fingerprint. The monitor receives the failure; the eval dataset or metrics path receives the outcome and cost. Later, the experiment report joins those two views by dimensions the application owns.
Keep the fingerprint conservative. Exception type plus operation is often a better starting point than the full message, because messages can contain request-specific IDs and split one defect into hundreds of groups. On the other hand, a fingerprint that uses only the exception type can merge unrelated failures. This is a trade-off, not a formatting detail.
I prefer the boring key.
The cost value also needs a precise meaning. In the example below it is the cost for the model call associated with this application operation, not an estimate for the whole request and not a vendor invoice total. That boundary prevents a retrying agent loop from quietly turning one “request cost” field into an ambiguous sum.
Build the attribution record before choosing a vendor
Start by proving that the monitoring account can return grouped errors. This minimal client uses the Python standard library, keeps the key in an environment variable, sets the HTTP method explicitly, honors Retry-After on a 429 response, and surfaces the real response body on other HTTP errors.
from __future__ import annotations
import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def list_error_groups(max_attempts: int = 4) -> object:
api_key = os.environ["INFRAI_API_KEY"]
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
request = Request(
f"{base_url}/errors/groups",
method="GET",
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
)
for attempt in range(max_attempts):
try:
with urlopen(request, timeout=15) as response:
return json.load(response)
except HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"Error API returned {exc.code}: {body}") from exc
retry_after = exc.headers.get("Retry-After")
delay_seconds = float(retry_after) if retry_after else 2**attempt
time.sleep(delay_seconds)
raise RuntimeError("Error group query exhausted its retry budget")
print(json.dumps(list_error_groups(), indent=2))
Run it with INFRAI_API_KEY and the service's versioned API base URL set in the process environment. The code deliberately makes one read-only call; capturing an exception needs the documented request schema, and copying that schema from public discovery is safer than guessing fields in an article.
The next runnable example is the app-owned half of the integration. It produces one success record and one grouped failure record, then compares variants inside a cohort. Replace the in-memory lists with your eval store and chosen error capture client once the shape is settled.
from __future__ import annotations
from collections import defaultdict
from dataclasses import asdict, dataclass
from decimal import Decimal
import hashlib
import json
from typing import Callable
@dataclass(frozen=True)
class RunContext:
tenant_cohort: str
experiment_variant: str
operation: str
@dataclass(frozen=True)
class RunResult:
context: RunContext
model_cost_usd: Decimal
ok: bool
error_fingerprint: str | None
def fingerprint(operation: str, exc: Exception) -> str:
stable_input = f"{operation}:{type(exc).__name__}"
return hashlib.sha256(stable_input.encode("utf-8")).hexdigest()[:16]
def run_once(
context: RunContext,
model_cost_usd: Decimal,
workload: Callable[[], str],
) -> RunResult:
try:
workload()
return RunResult(context, model_cost_usd, True, None)
except Exception as exc:
return RunResult(
context,
model_cost_usd,
False,
fingerprint(context.operation, exc),
)
def summarize(results: list[RunResult]) -> list[dict[str, str | int]]:
totals: dict[tuple[str, str], dict[str, Decimal | int]] = defaultdict(
lambda: {"runs": 0, "failures": 0, "cost_usd": Decimal("0")}
)
for result in results:
key = (result.context.tenant_cohort, result.context.experiment_variant)
totals[key]["runs"] += 1
totals[key]["failures"] += int(not result.ok)
totals[key]["cost_usd"] += result.model_cost_usd
return [
{
"tenant_cohort": cohort,
"experiment_variant": variant,
"runs": int(values["runs"]),
"failures": int(values["failures"]),
"cost_usd": str(values["cost_usd"]),
}
for (cohort, variant), values in sorted(totals.items())
]
def succeeds() -> str:
return "accepted"
def fails() -> str:
raise TimeoutError("model response exceeded the application deadline")
results = [
run_once(
RunContext("mid_market", "control", "draft_article"),
Decimal("0.018"),
succeeds,
),
run_once(
RunContext("mid_market", "candidate", "draft_article"),
Decimal("0.024"),
fails,
),
]
print(json.dumps(summarize(results), indent=2))
print(json.dumps(asdict(results[1]), indent=2, default=str))
Two records prove nothing about an experiment; they only make the mechanics inspectable. In production, the decision rule should live in the eval harness: compare quality, failure rate, and total model cost for each cohort and variant over a declared sample. Do not optimize cost by discarding failed calls. A failed call can still incur model cost, and excluding it rewards the least reliable variant. If one variant makes three model attempts before failing while another succeeds on its first attempt, charging only successful runs makes the retry-heavy path look artificially efficient. The eval must count known call costs before it judges the winner.
This is the notebook-to-production handoff I care about most. The notebook can explore groupings freely. Production needs a versioned context contract, bounded labels, deliberate redaction, and a test that confirms every experiment path emits the same dimensions.
Where does the simplest option stop being simple?
An errors API is attractive when a junior developer needs four basic actions: send exceptions, inspect grouped events, search recent failures, and mark groups resolved. Infrai fits that narrow requirement and gives a team one key for everything, one bill, and one plain REST API instead of a separate SDK for each supported capability. The breadth behind that simple contract is 295 routes across 20 modules. Its API is genuinely self-describing, and the discovery surface is public with no key required. Its practical boundary is alert delivery: there are no native threshold rules or notification routes, so alerts require polling and your own delivery logic.
That missing alert layer is not cosmetic. Polling has to remember a cursor, avoid duplicate pages, survive rate limits, decide when repeated events constitute an incident, and route the result somewhere people will see it. For a low-volume internal tool, that may remain small. For a customer-facing SaaS with an on-call expectation, it becomes a service you own.
The same boundary appears in debugging. The basic API does not provide distributed-trace queries or span trees, source-map deobfuscation, crash symbolication, Electron minidump parsing, or Session Replay. Logs may carry trace_id and span_id for correlation, but those fields do not create a trace explorer. Silent scheduled-job failures are another category entirely; pair error capture with a heartbeat service such as Healthchecks because “the task never ran” produces no exception to capture.
Be strict here. My first instinct is to keep the stack small, but the trade-off changes once someone carries a pager. I would accept the basic API only when the team deliberately owns that polling loop; otherwise, built-in routing wins. If the team says it needs alerts “later,” ask who will build and operate them, and put that work beside the vendor subscription in the decision record.
A fair comparison for a small SaaS
The products overlap, but they are not interchangeable. This table focuses on the capabilities that change the implementation, not a temporary price sheet.
| Option | Best fit in this scenario | Engineering boundary |
|---|---|---|
| Plain errors API | Low-friction capture, grouped search, and resolution inside an existing app stack | You own notification polling and routing; no trace tree, source-map processing, symbolication, replay, or heartbeat monitoring |
| Sentry | Teams that need a dedicated production-debugging workflow and configurable event grouping | More observability surface than a team seeking only a searchable exception inbox may need |
| Rollbar | Teams that prefer a dedicated error-monitoring product with built-in notifications and richer debugging | Cost attribution for an AI cohort experiment still belongs in application instrumentation and the eval pipeline |
| Datadog | Teams evaluating error tracking as part of a broader observability purchase | A broad platform evaluation is a different project from adding a small exception inbox |
| Grafana | Teams already assembling an observability stack around shared telemetry | The team must decide which components own error grouping and notification delivery |
| Better Stack | Teams comparing an integrated operational monitoring option | Experiment cost and cohort semantics still remain application-owned |
| Healthchecks | Detecting cron and worker jobs that fail by never running | It complements exception monitoring; it is not the grouped error store |
Sentry documents how stack traces, exception data, messages, and fingerprints influence event grouping. That is directly relevant when tenant-specific values threaten to fragment one underlying defect. Rollbar belongs on the shortlist when built-in notifications and richer production debugging are requirements. Neither choice removes the need to own experiment semantics: candidate has meaning to your release process, and a monitoring vendor should not become the source of truth for it.
There is also a privacy boundary. Before attaching tenant context, use cohort labels rather than customer names, exclude prompts and generated content by default, and decide how an individual user's data can be found and deleted. A logging path without a per-user deletion interface cannot, by itself, satisfy a workflow that depends on targeted erasure. Retention and export requirements deserve the same review before adoption.
No prompt belongs there.
Pricing should come after fit. Compare the billing model against expected event volume and data retention, but avoid treating today's unit price as architecture; alert ownership and missing debugging context can dominate the real operating burden.
The production checklist is a set of tests
Before launch, I would turn the integration review into executable checks. Force one known exception in each experiment variant and assert that its cohort, operation, and fingerprint arrive intact. Repeat the exception with a different request ID and verify that it joins the same group. Then raise a different exception in the same operation and confirm it does not.
Test redaction with an email address and a prompt-shaped payload. Test retry behavior under rate limiting, and make the capture path non-blocking with a bounded failure policy so the monitor cannot take down the product it observes. Verify that a failed model call still contributes its known cost to the experiment total. Tiny omissions here skew the evaluation.
Finally, rehearse the quiet failures. Stop a scheduled worker before it begins and confirm the heartbeat monitor reports the missed run. Trigger enough repeated errors to cross your operational threshold and confirm a human receives the notification. If the selected errors API has no native routing, that rehearsal must cover the poller, deduplication, and delivery path you built. Document who owns it.
The decision is now straightforward: use the lean capture surface when searchable grouped exceptions are the whole requirement and the team accepts the operational work around alerts. Use Sentry or Rollbar when on-call routing and richer debugging are part of “done.” In both cases, keep cohort and model-cost attribution in the application contract. That is what lets an eval answer the business question instead of merely producing a cleaner stack trace.
Top comments (0)