If you send user text to an LLM API from India, there's a good chance it contains an Aadhaar
number, a PAN, a UPI ID or a phone number. India's DPDP Act sets a compliance deadline of
13 May 2027, and every prompt that leaves your servers is a data-sharing decision.
The usual answer is "run it through Presidio first". Presidio is excellent, mature, and the
right default for generic PII. But it was built around US/EU data. So we measured how it does
on Indian identifiers.
The benchmark
IndiaPII-Bench v1.0 is 2,000 synthetic documents with 13,468 labelled spans across 17 Indian
entity types. Every identifier is synthetic but valid: Aadhaar numbers pass the Verhoeff checksum,
GSTINs pass their check digit, and so on. The corpus is published on Hugging Face
(maskflow-ai/indiapii-bench), and the harness is one command, so anyone can reproduce it.
We compared four tools: stock Presidio, Presidio with custom Aadhaar + PAN recognizers,
mask-privacy, and MaskFlow (the open-source library we built).
What we found
| Entity | MaskFlow | Presidio (stock) | Presidio + custom | mask-privacy |
|---|---|---|---|---|
| Aadhaar | 98.4% | — | 96.6% | — |
| PAN | 100% | — | 100% | — |
| GSTIN, IFSC, UPI VPA | 100% | — | — | — |
| Indian mobile | 99.0% | 94.9% | 94.9% | 42.2% |
| Person name | 52.4% | 30.4% | 30.4% | 30.4% |
| Indian address | 43.3% | 48.2% | 48.2% | 50.7% |
(Partial-overlap F1. "—" means the tool made no predictions for that type.)
Three takeaways:
- Default Presidio doesn't detect Aadhaar, PAN, GSTIN, IFSC or UPI IDs. You can add recognizers, and two good regexes get Aadhaar to 96.6%. But most teams don't know they need to.
- Checksums matter. A naive 12-digit regex scores 54.4% on Aadhaar, because order numbers and account numbers look the same. Validating the Verhoeff check digit is what gets you to 98%+.
- Names and addresses are hard for everyone. We lead on names and lose on addresses, where mask-privacy is ahead. We publish both.
What MaskFlow does
MaskFlow replaces PII with typed, reversible placeholders before text leaves your process, and
restores the real values in the response:
from maskflow import mask, unmask
# Synthetic Aadhaar number: passes the Verhoeff checksum, belongs to no one.
result = mask("My Aadhaar is 2346 8907 6549 and you can reach me at alice@example.com.")
result.masked_text
# "My Aadhaar is <AADHAAR_1> and you can reach me at <EMAIL_1>."
unmask(result.masked_text, result.mapping) # original text, restored
The model sees <AADHAAR_1>, so it can still write a sensible reply, and your user sees their
real details. It's MIT-licensed, runs entirely on your infrastructure, and plugs into LiteLLM,
LangChain, LlamaIndex and MCP, or runs as an OpenAI-compatible gateway.
Honest limits
- Names in Devanagari script are detected only after a cue such as
नाम:orश्री. - Address detection is weaker than we'd like (43.3% F1).
- The benchmark is synthetic. Real text is messier, so run
maskflow bench --my-dataon your own labelled data.
Code: https://github.com/maskflow/maskflow · Benchmark: https://maskflow.in/benchmark
If it's useful, a ⭐ on GitHub helps other teams find it.
Top comments (0)