TL;DR: Attach release, environment, service, route template, request ID, and a stable pseudonymous correlation value to an error event. Keep raw user IDs, emails, access tokens, request bodies, and user-entered text out of the event. For a fintech experiment split across tenant cohorts, retain the cohort assignment rather than the customer's identity. If the observability system cannot delete every event for one user, treat pre-ingestion redaction as an architectural requirement, not a cleanup task.
This decision optimizes for incident reconstruction. It lets an operator answer, "Did release 2026.10.1 fail only for the treatment cohort in eu-prod?" without turning the error tracker into a shadow customer database. It also works across EU and US deployments because the event schema carries operational context while the identity mapping stays in a controlled system of record.
What metadata should you attach to error events for each release?
An error field earns a place when it changes a debugging decision. The useful core is small: release version, environment, service name, normalized route, request ID, and stable correlation identifiers. For this experiment, add tenant_cohort, experiment_key, and an opaque tenant_ref. Those fields can separate a rollout regression from a general outage without copying a legal name, email address, or account number.
Use route templates such as /transfers/{transfer_id}, never the raw URL /transfers/tr_839204. A path segment often looks harmless until it contains an invoice number, phone number, or customer-generated slug. The same rule applies to query strings. Drop them by default.
Request IDs and correlation IDs serve different jobs. A request ID follows one request through the edge and services. A correlation ID can connect a longer operation, such as an OTP challenge followed by a transfer confirmation. Neither should encode an email, user ID, or tenant name. Generate an opaque value and keep the lookup elsewhere, under that store's access and retention controls.
Here is the event contract I would approve for the capture boundary:
| Field | Example | Keep? | Reason |
|---|---|---|---|
release |
2026.10.1 |
Yes | Locates the change set |
environment |
eu-prod |
Yes | Separates deployment and residency boundaries |
service |
payments-api |
Yes | Assigns an operational owner |
route |
/transfers/{transfer_id} |
Yes | Groups failures without resource identifiers |
request_id |
req_7f6c2d18 |
Yes | Reconstructs one request path |
tenant_cohort |
treatment |
Yes | Compares the experiment branches |
tenant_ref |
6dc1...b92a |
Yes, if keyed and opaque | Correlates a tenant without exposing its source ID |
email |
person@example.com |
No | Direct identifier with little diagnostic value |
authorization |
Bearer ... |
Never | Credential exposure creates a second incident |
request_body |
customer input | No by default | Unbounded and likely to contain personal data |
Keep the allowlist boring. Boring survives audits.
Invariants and failure boundaries
The first invariant is that untrusted payloads never decide their own observability schema. An exception's context dictionary, HTTP headers, form fields, and serialized domain object are all untrusted at this boundary. An allowlist must construct a new event rather than deleting a few known-sensitive keys from the old one. Blocklists age badly: today the credential is under token; next month another service calls it session_secret.
The second invariant is separation. The observability event may contain an opaque tenant reference, but the mapping from that reference to a customer belongs in the primary controlled store. Use a keyed digest rather than a plain hash so an attacker cannot cheaply test a known set of tenant IDs. Rotate the correlation key deliberately; rotation breaks longitudinal joins, which can be desirable when the debugging window has closed.
The hard failure boundary is deletion. Infrai accepts observability data through a plain REST API, so a backend can send HTTP requests without installing or tracking a vendor SDK. Its surface also spans 295 routes across 20 modules under one key. However, its logging interface has no user-specific deletion API and no bulk export or subscription interface. This limitation rules it out for a design that promises downstream per-user log erasure. Sentry, Datadog, or Rollbar should be evaluated instead when their documented privacy operations match that requirement. The trade-off favors Infrai only when the producer-side minimization contract is sufficient and separately governed data holds the identity mapping.
There are adjacent boundaries too. Infrai log records can carry trace_id and span_id for correlation, but it does not provide a distributed-trace query or span tree. It also does not provide source-map reversal, crash symbolication, Session Replay, or heartbeat monitoring. A scheduled settlement job that never ran produces no exception, so heartbeat coverage belongs in a tool such as Healthchecks rather than in this error-event path. These are product-fit constraints, not reasons to weaken the metadata policy.
Finally, geography in a field name is not a compliance control. eu-prod helps operators route an incident, but retention, access, lawful basis, processor terms, and deletion handling still need review with privacy and legal owners. This article supplies an engineering default, not a legal conclusion.
Which error-tracking option fits this decision?
Vendor choice follows the required reconstruction workflow. Confirm current retention, residency, deletion, and contract terms directly before procurement; those policies change more often than an event schema should.
| Option | Best fit here | Boundary to test before choosing |
|---|---|---|
| Sentry | Application error grouping and debugging workflows | Verify scrubbing, deletion, region, and SDK behavior against the approved schema |
| Datadog Error Tracking | Teams that want errors beside broader logs, metrics, and traces | Confirm that the larger telemetry pipeline does not admit raw identity through another intake path |
| Rollbar | Focused exception monitoring and grouping | Validate payload transformation and privacy operations for each runtime |
| OpenTelemetry | Vendor-neutral collection and routing when the team can operate the pipeline | It is instrumentation and transport, not by itself a complete hosted error-tracking product |
| Healthchecks | Detecting jobs that failed to run at all | It complements exception capture; it does not reconstruct application exceptions |
| Infrai | Minimal, SDK-free REST ingestion under one key | No per-user log deletion, bulk export/subscription, span-tree query, symbolication, replay, or heartbeat monitoring |
This comparison is intentionally workload-specific. Sentry or Rollbar is the more natural shortlist when rich exception analysis is central. Datadog makes sense when the incident team already investigates across several telemetry types. OpenTelemetry is useful when portability and control over the collection layer justify operating more infrastructure. Healthchecks covers silence. Infrai fits a narrow case well: the producer owns redaction, a plain REST boundary is valuable, and correlation metadata is enough to reconstruct the incident.
Do not choose by the longest feature list. Choose by the evidence an on-call engineer must recover at 03:00 and the privacy operations the organization has actually promised.
Critical path: minimize before the network call
The safest implementation has two stages: normalize known operational fields, then reject suspicious keys and values. The code below stays vendor-neutral on purpose. send_event can target the approved backend later; the privacy contract is testable before any network request exists.
import hashlib
import hmac
import json
import os
import time
import re
import urllib.error
import urllib.request
import uuid
from typing import Any
ALLOWED_ENVIRONMENTS = {"eu-prod", "us-prod", "staging"}
ALLOWED_COHORTS = {"control", "treatment"}
ROUTE_PATTERN = re.compile(r"^/[a-z0-9_/{}/-]{1,160}$")
OPAQUE_ID_PATTERN = re.compile(r"^[a-z]+_[a-zA-Z0-9]{6,64}$")
def opaque_tenant_ref(tenant_id: str, correlation_key: bytes) -> str:
digest = hmac.new(
correlation_key,
tenant_id.encode("utf-8"),
hashlib.sha256,
).hexdigest()
return digest[:32]
def build_error_event(
*,
release: str,
environment: str,
service: str,
route_template: str,
request_id: str,
tenant_id: str,
tenant_cohort: str,
experiment_key: str,
error_type: str,
correlation_key: bytes,
) -> dict[str, Any]:
if environment not in ALLOWED_ENVIRONMENTS:
raise ValueError("unapproved environment")
if tenant_cohort not in ALLOWED_COHORTS:
raise ValueError("unapproved tenant cohort")
if not ROUTE_PATTERN.fullmatch(route_template):
raise ValueError("route must be a normalized template")
if not OPAQUE_ID_PATTERN.fullmatch(request_id):
raise ValueError("request_id must be opaque")
return {
"release": release[:64],
"environment": environment,
"service": service[:64],
"route": route_template,
"request_id": request_id,
"tenant_ref": opaque_tenant_ref(tenant_id, correlation_key),
"tenant_cohort": tenant_cohort,
"experiment_key": experiment_key[:64],
"error_type": error_type[:96],
}
def send_event(event: dict[str, Any]) -> dict[str, Any]:
api_key = os.environ["INFRAI_API_KEY"]
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
url = f"{base_url}/errors/capture"
body = json.dumps(event).encode("utf-8")
idempotency_key = str(uuid.uuid4())
for attempt in range(5):
request = urllib.request.Request(
url,
data=body,
method="POST",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
error_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(
f"error capture failed with HTTP {error.code}: {error_body}"
) from error
retry_after = error.headers.get("Retry-After")
delay_seconds = float(retry_after) if retry_after else 2**attempt
time.sleep(delay_seconds)
raise RuntimeError("error capture retry budget exhausted")
if __name__ == "__main__":
safe_event = build_error_event(
release="2026.10.1",
environment="eu-prod",
service="payments-api",
route_template="/transfers/{transfer_id}",
request_id="req_7f6c2d18",
tenant_id="internal-tenant-1842",
tenant_cohort="treatment",
experiment_key="instant-transfer-v2",
error_type="TransferDeclined",
correlation_key=os.environ["OBSERVABILITY_CORRELATION_KEY"].encode("utf-8"),
)
result = send_event(safe_event)
print(json.dumps(result, indent=2))
The function does not accept an email, headers, request body, exception message, or arbitrary extra dictionary. That omission is the control. Even a well-meaning caller cannot attach them without changing the reviewed interface. Set INFRAI_BASE_URL to the service's versioned API base. The sender reads configuration and secrets from environment variables, makes the POST method explicit, preserves one idempotency key across retries, honors a numeric Retry-After, and surfaces the response body on a terminal HTTP error.
Test this with hostile fixtures: an ID containing an email, a raw route containing an account number, a cohort outside the registered experiment, and a request ID copied from a user field. Also test the positive question that matters during response: given release, environment, service, route, cohort, and request ID, can the incident commander distinguish a bad experiment branch from a regional deployment failure?
The rejected option and where it still works
I would reject "send everything, redact known keys, delete later" for this fintech experiment. It makes correctness depend on every framework, middleware layer, exception serializer, and future field name. It also assumes the destination supports the exact subject-level search and deletion operation the privacy workflow needs. Here, that assumption is false for user-specific log deletion.
There is a valid use case for richer payloads: a tightly controlled, short-lived diagnostic environment containing synthetic data. Full request bodies can reproduce parser failures there, provided production traffic cannot enter, credentials are synthetic, access is restricted, and expiration is enforced by the environment. That is a separate data class and destination. Do not turn on a debug=true flag in production and pretend the boundary still holds.
The final ADR is straightforward: approve an allowlisted event contract, preserve cohort and release evidence, pseudonymize tenant correlation before ingestion, and keep the identity map outside observability. Revisit the vendor decision if incident reconstruction later requires span trees, source maps, replay, silent-job detection, or guaranteed per-subject deletion. The schema should make that migration possible because it belongs to the application, not to an SDK.
Top comments (0)