A recent incident showed that redacted reports still revealed Google Data Center water and electricity usage. The root cause was a careless redaction script that only performed simple string replacements, leaving numeric patterns untouched. This article walks you through the exact mistake, a working fix, and the tradeoffs you’ll face when securing sensitive infrastructure metrics.
What you'll learn
- Why naive
str.replacebreaks redaction. - How to build a regex that catches usage numbers in multiple formats.
- Tradeoffs between speed, accuracy, and maintainability.
- Common failure modes and simple tests to catch them.
Common Mistake: Over‑Simplistic String Replace
Developers often reach for str.replace when they need to hide numbers. It works for exact matches but fails when the same value appears in different units or with extra whitespace. A report containing "12,500 MWh" and "8,300 gallons" can slip through because the script only looks for the raw digits.
## naive redaction – leaves numbers untouched
def naive_redact(text: str) -> str:
# replace known labels with placeholders
text = text.replace("MWh", "[REDACTED]")
text = text.replace("gallons", "[REDACTED]")
return text
This snippet shows the intuitive but flawed approach. It only masks the unit labels, not the numeric values, so the underlying usage remains visible.
Robust Redaction with Regex
A single regex can capture both the number and its unit, allowing you to replace the whole token. The pattern r'\b\d{1,3}(?:,\d{3})*(?:\.\d+)?\s*(?:MWh|gallons)\b' matches numbers with optional thousands separators and optional decimal parts, followed by a unit.
import re
usage_pattern = re.compile(r'\b\d{1,3}(?:,\d{3})*(?:\.\d+)?\s*(?:MWh|gallons)\b')
def robust_redact(text: str) -> str:
return usage_pattern.sub('[REDACTED]', text)
Why this works: the regex is anchored to word boundaries, preventing partial matches, and it captures the whole usage token in one go.
Tradeoffs: Performance vs Coverage
| Approach | Coverage | Performance | When to Use |
|---|---|---|---|
| Simple replace | Low | Fast | Quick scripts, low‑risk data |
| Regex | High | Moderate | Most production redaction jobs |
| Library (e.g., Presidio) | Very High | Slow | Complex PII pipelines, compliance‑heavy contexts |
Choose regex when you need reliable masking without the overhead of a full‑blown NLP library.
Failure Modes and Testing
- Overlapping patterns – a number may match both a unit and a surrounding label. Use non‑capturing groups and order patterns from most specific to least.
- False positives – the regex could match numbers that are not usage metrics (e.g., timestamps). Add context checks or whitelist known units.
- Unicode and locale – thousands separators vary by region. Adjust the pattern to accept locale‑specific separators if needed.
Test by feeding sample redacted documents into a small script that asserts no usage numbers remain after processing.
Key Takeaways
- Never rely on simple string replace for numeric redaction; it leaves the core data exposed.
- A well‑crafted regex can capture both the value and its unit, providing strong coverage.
- Understand the performance impact; regex is a good middle ground between speed and accuracy.
- Anticipate overlapping matches and false positives, and write unit tests that verify the redacted output.
- For enterprise‑grade compliance, consider dedicated libraries, but start with regex for most internal pipelines.
Source
Improper redaction reveals Google Data Center water and electricity usage
I added concrete code examples, a comparison table, and failure‑mode analysis that the original article omitted.
Support this work
These write-ups are researched and published with no paywall, sponsor, or tracking. If one saved you an afternoon, a small tip keeps them coming.
USDT, USDC or USDD · TRC-20 (Tron)
TFTNsfyomKrnUutRjBTGVULp19ByW29KbY
Top comments (0)