DEV Community

sodanisomya
sodanisomya

Posted on

Out of the box, Presidio doesn't catch Aadhaar numbers. We measured.

If you send user text to an LLM API from India, there's a good chance it contains an Aadhaar
number, a PAN, a UPI ID or a phone number. India's DPDP Act sets a compliance deadline of
13 May 2027, and every prompt that leaves your servers is a data-sharing decision.

The usual answer is "run it through Presidio first". Presidio is excellent, mature, and the
right default for generic PII. But it was built around US/EU data. So we measured how it does
on Indian identifiers.

The benchmark

IndiaPII-Bench v1.0 is 2,000 synthetic documents with 13,468 labelled spans across 17 Indian
entity types. Every identifier is synthetic but valid: Aadhaar numbers pass the Verhoeff checksum,
GSTINs pass their check digit, and so on. The corpus is published on Hugging Face
(maskflow-ai/indiapii-bench), and the harness is one command, so anyone can reproduce it.

We compared four tools: stock Presidio, Presidio with custom Aadhaar + PAN recognizers,
mask-privacy, and MaskFlow (the open-source library we built).

What we found

Entity MaskFlow Presidio (stock) Presidio + custom mask-privacy
Aadhaar 98.4% — 96.6% —
PAN 100% — 100% —
GSTIN, IFSC, UPI VPA 100% — — —
Indian mobile 99.0% 94.9% 94.9% 42.2%
Person name 52.4% 30.4% 30.4% 30.4%
Indian address 43.3% 48.2% 48.2% 50.7%

(Partial-overlap F1. "—" means the tool made no predictions for that type.)

Three takeaways:

  1. Default Presidio doesn't detect Aadhaar, PAN, GSTIN, IFSC or UPI IDs. You can add recognizers, and two good regexes get Aadhaar to 96.6%. But most teams don't know they need to.
  2. Checksums matter. A naive 12-digit regex scores 54.4% on Aadhaar, because order numbers and account numbers look the same. Validating the Verhoeff check digit is what gets you to 98%+.
  3. Names and addresses are hard for everyone. We lead on names and lose on addresses, where mask-privacy is ahead. We publish both.

What MaskFlow does

MaskFlow replaces PII with typed, reversible placeholders before text leaves your process, and
restores the real values in the response:

from maskflow import mask, unmask

# Synthetic Aadhaar number: passes the Verhoeff checksum, belongs to no one.
result = mask("My Aadhaar is 2346 8907 6549 and you can reach me at alice@example.com.")
result.masked_text
# "My Aadhaar is <AADHAAR_1> and you can reach me at <EMAIL_1>."
unmask(result.masked_text, result.mapping)  # original text, restored
Enter fullscreen mode Exit fullscreen mode

The model sees <AADHAAR_1>, so it can still write a sensible reply, and your user sees their
real details. It's MIT-licensed, runs entirely on your infrastructure, and plugs into LiteLLM,
LangChain, LlamaIndex and MCP, or runs as an OpenAI-compatible gateway.

Honest limits

  • Names in Devanagari script are detected only after a cue such as नाम: or श्री.
  • Address detection is weaker than we'd like (43.3% F1).
  • The benchmark is synthetic. Real text is messier, so run maskflow bench --my-data on your own labelled data.

Code: https://github.com/maskflow/maskflow · Benchmark: https://maskflow.in/benchmark

If it's useful, a ⭐ on GitHub helps other teams find it.

Top comments (0)