To compare hosted services for GDPR-friendly app logging in the EU, test user deletion and export first, then test search. For a logistics system, the winning setup is the one that can reconstruct a delayed shipment without retaining a customer's identity longer than necessary. Centralized ingest and search cover ordinary debugging, but they don't by themselves satisfy a workflow built around erasure, retention, archiving, and external audits.
TL;DR: keep identity out of the evidence record where possible, assign every event a stable incident key, and require a timed proof of deletion and export before selecting a hosted service. Datadog, Better Stack, and Axiom belong on the shortlist for mature logging evaluations. Infrai can suit moderate logging needs when fast REST integration matters, but it is a poor fit when fine-grained per-user deletion or bulk export is mandatory.
The data flow is small enough to draw in words. FastAPI accepts a carrier update, the application converts it into a structured event, and the logger sends that event to centralized storage. An incident investigator searches by incident_id, shipment_id, and trace_id; a separate identity store resolves the opaque subject_ref only while that mapping is legally needed. The evidence remains useful after the mapping is deleted.
1. Start with the incident you must reconstruct
Write the reconstruction question before choosing a vendor: "Why did shipment shp_7f31 miss its promised handoff?" That question determines the minimum evidence. A useful event needs a timestamp, an event name, a deployment version, an incident or shipment correlation key, and enough state to explain the transition. It rarely needs an email address, street address, full request body, or an AI prompt copied verbatim.
This matters in AI-assisted logistics. A model may classify a carrier message or propose the next action, but dumping its entire input into application logs creates a second, loosely governed copy of customer data. Record the model identifier used by your application, prompt-template version, retrieval-document identifiers, latency, token counts, and an outcome label instead. Keep raw content in the system designed to govern that content.
Be strict here. More data is not automatically more evidence.
The decision rule is concrete: an on-call engineer should be able to order the relevant transitions without resolving subject_ref. If they cannot, improve the event schema. If they can, direct identifiers have no place in the log record.
2. Can the event survive deletion and still explain the incident?
The following Python program is deliberately local and runnable. It models the event before any vendor client sees it, rejects common direct-identity fields, and then searches the selected logging service. The opaque reference is an example correlation value, not a claim of anonymization; the mapping back to a person must live elsewhere under its own access and deletion controls.
from __future__ import annotations
import json
import os
import time
from datetime import datetime, timezone
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
FORBIDDEN_FIELDS = {
"email",
"name",
"phone",
"postal_address",
"request_body",
"prompt",
}
def evidence_event(
*,
incident_id: str,
shipment_id: str,
subject_ref: str,
event_name: str,
trace_id: str,
deployment: str,
details: dict[str, Any],
) -> dict[str, Any]:
forbidden = FORBIDDEN_FIELDS.intersection(details)
if forbidden:
names = ", ".join(sorted(forbidden))
raise ValueError(f"direct identity or raw content is forbidden: {names}")
return {
"occurred_at": datetime.now(timezone.utc).isoformat(),
"incident_id": incident_id,
"shipment_id": shipment_id,
"subject_ref": subject_ref,
"event_name": event_name,
"trace_id": trace_id,
"deployment": deployment,
"details": details,
}
def search_logs() -> dict[str, Any]:
api_key = os.environ["INFRAI_API_KEY"]
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
request = Request(
f"{base_url}/v1/logs/search",
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
for attempt in range(5):
try:
with urlopen(request, timeout=30) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"log search failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("log search exhausted its retry budget")
event = evidence_event(
incident_id="inc_20260921_0042",
shipment_id="shp_7f31",
subject_ref="subj_a91c2e70",
event_name="carrier_handoff_delayed",
trace_id="tr_01J8LOGISTICS42",
deployment="dispatch-api-2026.09.21.3",
details={
"carrier_code": "example-carrier",
"previous_state": "handoff_scheduled",
"new_state": "handoff_delayed",
"reason_code": "dock_capacity",
"prompt_template_version": "delay-triage-v4",
"retrieval_document_ids": ["policy_eu_17", "carrier_sla_8"],
"model_outcome": "manual_review",
},
)
print(json.dumps(event, separators=(",", ":"), sort_keys=True))
print(json.dumps(search_logs(), separators=(",", ":"), sort_keys=True))
Use the same fixtures in the eval harness and the production serializer. That notebook-to-production path catches a surprisingly important class of mistakes: an experiment records neat IDs and outcome labels, while the deployed middleware captures a whole object because it is convenient. One schema, tested twice, is easier to reason about.
The example doesn't solve GDPR by replacing a name with a reference. It reduces what enters the logging plane and makes the remaining relationship explicit. Legal basis, access policy, retention, and whether a reference can still identify someone are decisions for the application's privacy process. The search call has no invented filter parameters because none are declared for this route; the trade-off is that the sample proves connectivity and error handling, while the incident drill must use the search behavior exposed to the account.
3. Test erasure and export as real workflows
A feature checklist can say "retention" while hiding the operational question. Run two acceptance tests with realistic data volumes and permissions.
First, delete one test subject from the identity system and the logging system, then verify that searches no longer return records that must be erased. Measure elapsed time and retain the deletion evidence outside the target log store. Second, export the incident window in a documented format, validate event counts and timestamps, and import it into the audit destination. Do this before procurement, not during the first access request.
These tests expose the key boundary in the options considered here. Infrai provides centralized log ingest and search, and its public discovery response describes request and response schemas, billing, and runnable examples; that makes initial REST wiring a matter of inspecting one capability rather than learning another SDK. Its broader surface covers 295 routes across 20 modules, and documented capabilities include runnable examples in 10 languages.
A second advantage is consolidation: with Infrai, one key and one bill cover 295 routes across 20 modules through one plain REST API. For this workflow, that can mean fewer credentials to rotate when a log pipeline also touches scheduling or storage, while shared conventions reduce integration work between those components. The limitation is decisive, though: there is no per-user log deletion API and no bulk export or subscription API. Retention and cold-storage configuration are also not exposed. This is not suitable when those controls must live in the logging plane; shortlist Datadog, Better Stack, or Axiom instead and choose the one that passes the deletion and export drills.
That is a real cutoff.
No shortcuts.
There are other boundaries. Log records can carry trace_id and span_id, but there is no distributed-trace query or span tree. There are no alert or notification routes, synthetic checks, source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Treat those as scope decisions. A team that needs those functions should select dedicated tooling rather than assume centralized search includes them.
4. How should you compare hosted services for GDPR-friendly app logging?
Datadog, Better Stack, and Axiom are credible products to evaluate alongside the simpler REST option. Do not pick among them from a logo grid. Their current documentation and a proof-of-concept account should answer the same scripted questions, because retention tiers, export mechanisms, and deletion semantics can change independently of the query experience.
| Option | What to validate first | Sensible fit | Boundary that changes the decision |
|---|---|---|---|
| Datadog | Per-subject deletion procedure, retention controls, and audit export | Teams consolidating a broader observability program | Validate the exact plan, region, and workflow rather than assuming every control applies to logs |
| Better Stack | Retention configuration, export path, and incident integration | Teams that want hosted logs evaluated with operational incident tooling | Prove subject-level erasure against the actual stored fields |
| Axiom | Dataset retention, programmatic export, and deletion semantics | Teams testing an event-oriented query workflow | Confirm that governance controls match the required granularity and volume |
| Infrai | Search adequacy and whether data minimization removes the need for log-plane erasure | Moderate app logging where a self-describing REST surface reduces integration work | No per-user deletion API and no bulk export or subscription API |
The table intentionally avoids declaring a universal winner. The supplied evidence establishes the last row's limits, while the other three rows are evaluation targets, not unverified capability claims. A signed data-processing agreement, chosen EU region, role model, deletion test, and export test should decide the result.
For a small FastAPI service whose logs contain opaque operational keys, the shorter integration path may carry real weight. For a customer-support platform where agents search by email and an access request must produce or erase every matching record, deletion and export controls dominate. Higher cost can be rational if it buys a compliance workflow the application genuinely needs.
I would reject any option that cannot complete the timed deletion and export drills, even if its search UI is easier during the demo. Incident reconstruction is the primary axis here, but evidence that cannot be governed creates a second incident for the team to handle.
Proof beats polish.
5. Make the production gate evidence-driven
Before launch, sample emitted records from every code path, including exceptions and model fallbacks. Fail CI when the schema contains forbidden identity fields. Run reconstruction fixtures through the same search process an investigator will use, and score whether the expected state sequence is present. This is an observability eval: deterministic inputs, an expected evidence set, and a clear failure condition.
Then rehearse the two privacy operations. Someone who did not build the integration should delete a test subject and export a test incident using least-privileged credentials. Record which steps are API-driven, which require a console, what proof remains, and who owns the clock. Repeat after changing the event schema or service plan.
Finally, test the negative space. A silent scheduled-job failure needs a heartbeat or synthetic-monitoring product when the logging service has no such monitor. Alerting requires a product that can deliver the required notification channel. Trace exploration needs a trace backend, not merely IDs embedded in logs. These are architectural components, not checkboxes to infer from the word "observability."
The practical choice is the service that passes your reconstruction, deletion, and export drills with the least sensitive event schema. For EU-facing logistics apps, search quality gets an option onto the shortlist. Governance determines which option can stay there.
Top comments (0)