The Boundary Problem
OpenTelemetry documentation emphasizes that implementers are responsible for protecting sensitive data, noting that the system cannot determine what is sensitive in your specific context. A common misconception is that configuring the OpenTelemetry Collector to redact data at the ingestion point cleans up the entire pipeline. It does not. If an application writes a raw session token to a local disk buffer, a temporary file, or a third-party SDK before the collector receives the span, that copy persists independently of downstream processing. The collector only sees what is sent to it; it cannot reach back into the application's memory or local storage to delete what was already written.
Hypothetical Scenario
Consider a hypothetical e-commerce service that logs a user's credit card number to a local file for debugging purposes before sending telemetry to the collector. Even if the collector is configured with a redaction processor to strip the credit_card attribute from incoming spans, the local file still contains the raw number. If that file is later archived, backed up, or accessed by a developer, the sensitive data remains exposed. The collector's redaction is a post-processing step, not a prevention mechanism. As the OpenTelemetry docs state, the best way to prevent collection is to not collect the data in the first place, following the principle of data minimization.
Practical Implications
Teams must treat the application boundary as the primary control point. Relying solely on collector-side redaction creates a false sense of security. Developers should review instrumentation libraries to ensure they do not inadvertently capture PII, and implement data minimization at the source. The collector's processors, such as attribute or transform, are useful for managing data that has already been collected, but they cannot undo copies made earlier in the pipeline. For more details, see the OpenTelemetry security guidance.
Top comments (1)
The local-buffer example suggests a useful regression test: send a synthetic sensitive marker through the instrumentation while the collector is unavailable, then inspect the application's retry spool and local logs before restoring delivery. Checking only the final exported span could miss the earlier copy you describe.
I'd also keep the marker in an exception-path test, not only the successful request. Do you test those pre-collector copies separately from the collector's redaction output?