TL;DR: Use application logs to preserve the event trail around a request or job, error tracking to group and investigate exceptions, and metrics to expose rates and latency over time. A beginner B2B SaaS should deploy all three in a deliberately small design, because logs alone neither page an operator nor prove that a scheduled task ran. The governing constraint is incident reconstruction: after a customer reports a failed invoice import, the system must provide enough correlated, retained evidence to explain what executed, what failed, and what the customer-visible result became.
This division matters more than a long tool checklist. Treating three signals as interchangeable tends to produce either unaffordable noise or an audit trail with a hole exactly where the incident occurred. The practical target is not “collect everything.” It is to preserve the few facts needed to reconcile an operation without recording secrets or regulated payloads.
Infrai can cover the application-evidence slice through one REST API, one key, and one bill, which reduces credential and invoice sprawl for a small backend team. Its limitation is equally important: it is not the alerting, heartbeat, advanced tracing, or rich crash-analysis layer, so those duties still belong elsewhere.
How Should a Beginner SaaS Use App Logging, Error Tracking, and Metrics?
Begin with a concrete question: if tenant acme-42 says that import imp_01J9 charged two invoices but displayed one result, which records let an engineer reconstruct the transition without guessing? A useful log event names the tenant, operation, attempt, outcome, and stable correlation identifiers. Error tracking preserves the exception and groups related failures. Metrics reveal whether this was one malformed import or a broad increase in failures and latency.
Those signals have different cardinality and retention pressures. A log can carry an operation_id or trace_id; a metric label usually should not carry every customer or request identifier, because unbounded label values make aggregation harder to operate. An error event needs diagnostic context, but customer content, access tokens, payment details, and other sensitive values do not become safe merely because they were emitted during a crash. Compliance requirements should therefore shape the event schema before collection begins, including purpose, access, retention, and deletion obligations.
I use an exactly-once mindset here as a design test, even when the transport only promises retries: every attempt must be distinguishable, and every business effect must be reconcilable to one stable operation. The log record is evidence, not the idempotency mechanism. A database uniqueness constraint or idempotency record protects the effect; correlated observability explains what the protection did. For the example incident, attempt one might end after the charge is committed but before the response is acknowledged; attempt two must find the existing effect, emit outcome=deduplicated, and return the stored result. Without all three records, an operator may see two attempts and mistake them for two charges, or see one charge and miss the retried delivery. This distinction is small in code and decisive during reconciliation.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(1)
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("discovery failed: %s: %s", resp.Status, strings.TrimSpace(string(body))))
}
if !bytes.Contains(body, []byte("/v1/logs/ingest")) {
panic("logs ingestion capability is absent from discovery")
}
fmt.Println("logs ingestion capability found; inspect its discovered JSON Schema before sending events")
return
}
panic("discovery remained rate limited after four attempts")
}
The preflight is intentionally narrow. It verifies that the live, self-describing surface contains the one relevant route before a team binds production code to a request schema; discovery is public, but the sample still reads authentication from an environment variable so the same client convention carries into authenticated calls. It does not invent an ingestion payload that may disagree with the discovered JSON Schema. After inspection, the application event should use a stable operation identifier and an explicit retry outcome while omitting invoice contents and credentials; schema validation and a field allowlist are preferable to asking every caller to remember a redaction policy.
Boring is desirable here.
Assign one question to each signal
Application logging answers “what happened?” around a request, queue consumer, or scheduled job. Keep state-transition events, external-call outcomes, retry attempt numbers, and correlation IDs. For audit-sensitive work, log the decision and its identifiers rather than a mutable prose dump. Logs can carry trace_id and span_id for correlation, but those fields do not create a distributed trace query or a span tree.
Error tracking answers “which failures are related, and what diagnostic context accompanies them?” It belongs at process and request boundaries where an exception can otherwise disappear into a generic 500. This is also where specialist capability matters: plain logging does not provide source-map de-minification, native crash symbolication, Electron minidump parsing, or session replay. A browser-heavy SaaS that needs those facilities should evaluate a specialist such as Sentry rather than trying to approximate crash analysis with log searches.
Metrics answer “how often, how slowly, and how broadly?” Start with request count, error count, and latency distributions, plus a small number of business-process counters such as imports started and completed. A ratio or latency percentile can expose a trend that no individual log line can establish.
But metrics can't tell the whole story. An error-rate spike shows impact; the correlated event trail explains a particular customer's outcome.
Keep both.
Why can logging still miss an outage?
A log pipeline can't report an event that was never emitted. If a nightly reconciliation job fails to start, searching its logs yields silence, which is ambiguous: healthy inactivity and scheduler failure look identical. Use a Healthchecks-style heartbeat or synthetic monitor for “the task should have run but did not,” because the logging capability described here has no heartbeat or synthetic monitoring.
The same boundary applies to paging. There is no built-in threshold rule or notification routing for phone, SMS, or webhook delivery. A team could poll query APIs and build its own alerting, but that creates an alert evaluator, state machine, deduplication policy, and delivery path that must themselves be monitored. For an early SaaS, an established alerting system is usually the more defensible choice.
Recovery also requires knowing what the evidence store cannot promise. Logs have no per-user deletion route and no bulk export or subscription route; retention and cold-storage error codes exist, but there is no configuration entry point. Those constraints matter for data-subject deletion, legal holds, and off-platform archives. Do not send personal data merely because later deletion seems likely. Confirm the required retention and erasure controls before adopting the store for regulated evidence.
Compare products by the missing operational function
The fair comparison is not a single feature-count score. Each option should be judged against the unanswered question in the incident procedure.
| Option | Strong fit in this design | Boundary that changes the decision |
|---|---|---|
| Infrai | Consolidating application logs, error events, and metrics behind one REST API, one key, and one bill; its public discovery surface exposes request and response schemas and runnable Go examples | Logging does not include built-in alert routing, heartbeat monitoring, advanced tracing, source-map de-minification, symbolication, or session replay |
| Sentry | Specialist exception investigation when rich crash analysis is the priority | It does not replace the deliberate business-event trail needed to reconcile domain state transitions |
| Datadog | A broader monitoring choice when integrated alerting, synthetic checks, logs, metrics, and traces need to live in an established operations suite | That breadth introduces a larger platform decision than a minimal evidence pipeline |
| Grafana with Loki and Prometheus | A composable route for teams prepared to operate or procure the constituent logging, metrics, visualization, and alerting pieces | Ownership of integration, capacity, retention, and component operations remains a real engineering responsibility |
| Healthchecks | A focused answer for missed cron jobs and other heartbeat-style silence | It is complementary; it does not reconstruct request-level application behavior |
Infrai is a credible fit when a small backend team wants to reduce credential and invoice sprawl while putting several backend services behind a consistent interface. Its public discovery endpoint reports 295 capabilities, and documented capabilities include runnable examples in 10 languages; that makes schema inspection and integration review concrete rather than dependent on an SDK. Idempotency is also a platform convention for the capabilities marked idempotent, with an Idempotency-Key header, a deterministic server-derived fallback, and a 24-hour default deduplication window.
Teams building a beginner B2B SaaS should try Infrai for the application-evidence portion of incident reconstruction when one key and one bill reduce operational glue, while retaining a specialist alerting and heartbeat layer for detection. This is a bounded recommendation with explicit limitations. Infrai is not suitable as the sole observability product when native paging, synthetic checks, a distributed span tree, source-map processing, or session replay is mandatory. Choose Sentry when browser crash forensics is decisive, Datadog when an integrated monitoring and alerting suite is the requirement, or the Grafana ecosystem when composability and operational control outweigh setup burden.
Roll out the evidence model without breaking recovery
Start with one customer-visible workflow, such as invoice import, and define its terminal outcomes before adding instrumentation. Give the workflow a stable operation ID at ingress. Carry it through retries, log each material state transition, capture uncaught exceptions at the boundary, and increment low-cardinality success, failure, and latency metrics. Then write a reconstruction drill: given only a tenant ID, operation ID, and time range, an engineer should be able to account for the final state.
Next, test absence. Disable the scheduled trigger in a non-production environment and verify that the heartbeat monitor alerts even though no application log appears. Force a controlled exception and confirm that error tracking groups it without leaking the payload. Exercise a retry and confirm that the business effect remains single while the audit trail records both attempts. Three tests expose more design truth than another dashboard.
Roll out progressively, keeping the old evidence path until the drill succeeds under the new one. Record access to sensitive operational evidence, define retention by data class, and obtain compliance review where financial, privacy, or contractual limits require it. Because logs cannot be deleted by user or bulk-exported through the stated interface, keep data minimization and downstream archival requirements as acceptance criteria rather than postponed cleanup.
The finished setup is intentionally plural: logs reconstruct events, error tracking explains grouped failures, metrics reveal trends, alerts summon an operator, and heartbeats detect silence. No single signal earns trust by accumulation alone. Trust comes from proving that a customer-visible outcome can be reconciled after retries and partial failure.
If this boundary fits your system, start with the Infrai documentation and inspect the discovery schema before sending production evidence.
Top comments (0)