DEV Community

Krish Verma
Krish Verma

Posted on

How I built a secret scrubber for my desktop AI assistant: 7 patterns, an entropy gate, and zero model calls

How I built a secret scrubber for my desktop AI assistant: 7 patterns, an entropy gate, and zero model calls

Ankita is my open-source desktop AI companion — "a little companion, a lot more possible". It lives on my desktop, runs scheduled jobs while I'm away, keeps transcripts, and fires Telegram alerts. All of that means one thing: the moment I paste a token into chat, that token wants to leak into session JSON on disk, daemon logs, alert messages, and memory files.

So I built a secret scrubber. One module — src/security/secret-scrubber.mjs, 116 lines — that every persistence boundary in the app calls before anything hits disk. And it never phones a model: the header comment says it plainly — "Persistence-boundary detection: compiled patterns, no network or model calls." Asking a model "is this a secret?" would mean sending the secret to a model. That defeats the point.

The three layers

1. Seven named patterns. PATTERNS is a small table of shapes everyone knows: AWS access keys (AKIA/ASIA plus 16 chars), GitHub tokens (ghp_, gho_, github_pat_, …), Slack tokens (xoxb-, xoxp-, …), Stripe live keys (sk_live_/rk_live_), PEM private-key blocks, Bearer … tokens, and connection-string passwords (postgres://alice:…@host/db, mongodb+srv://…). Shapes are cheap and precise, so they're the first pass.

2. Labeled tokens, judged by entropy. Real-world secrets don't always wear a badge. So the LABEL regex watches for api_key=, token:, secret "…", passwd=hunter2 — any label ending in the usual suspects. A labeled value only counts as a secret if it's at least 12 characters long and has Shannon entropy above 4.2 bits per character, unless the label is literally password/passwd/pwd. That single rule is what keeps "Use the token count to estimate cost" from being redacted while api_key=aB3cD4… gets caught. A CONTEXT_KEY regex does the same job for JSON object keys, so {"private_key": "…"} is scrubbed in nested objects, arrays, and even JSON strings.

3. Encoded wrappers, depth ≤ 2. A base64 blob hiding an AWS key is still an AWS key. Any candidate 16+ characters long gets URL-decoded and base64-decoded (bounded at 16,384 chars, max recursion depth 2), and if the decoded form contains a secret, the wrapper itself is flagged as encoded_secret and redacted in place. Data URIs (data:image/…) are exempt — decoding every pasted screenshot would be a performance and sanity disaster.

Overlapping hits are sorted by start and merged into single ranges, and redaction is idempotent: redact(redact(x)) === redact(x), because the marker format [REDACTED:type] is itself recognized and skipped on the next pass.

The scrub runs in two directions

Detection is only half the design. The module exports five functions — detect, redact, containsSecret, secretValues, and redactValue — and the last two matter most.

redactValue is the write path, and its doc comment is the whole philosophy: "A fresh copy for disk/log/export. Never mutate the live model's messages." Every persistence point — session transcripts, project files, profile memory, daemon alert messages — calls it by default. In src/core/sessions.mjs, saveSession(file, data, { transform = redactValue }): redaction is the default, not an opt-in.

SecretHistory (in the desktop process) handles the other direction. When you type "save this as my openai key", a save-intent regex plus a naming pattern sends the real value into the OS keychain under a slugified name (capped at 80 chars, with -1, -2 suffixes for multi-secret pastes). From then on, every copy of that value anywhere is swapped for [STORED:keychain:openai-key] — longest-first so partial overlaps can't corrupt a replacement, and existing markers are preserved. Raw values live only in memory or vault ciphertext. This is the echo protection: even if a secret lands in a draft, an export, or a pasted model reply, it gets replaced by its reference marker before it can spread.

Tested like an adversary

The test suite throws a 50-sample corpus at it: the shaped tokens, their URL-encoded forms, their base64 forms, and labeled fakes like passwd=hunter2. Every sample must be detected, every redaction must be idempotent, and no sample may survive cleaning. Then the false-positive gauntlet: "How do I reset my password?", "I need a secret birthday surprise.", "Use the token count to estimate cost." — zero of them may be touched, and typical messages must average below one millisecond. A secret detector has to be fast enough to run on every keystroke-adjacent boundary without anyone noticing.

The question I'm still arguing with myself about

redactValue deliberately never touches the live conversation. The running chat keeps the real secret, because the companion needs it to actually do its job — a scrubbed model can't call your API. Disk, logs, and exports get clean copies.

But is that the right boundary? The alternative is a vault-reference model: the live chat only ever sees [STORED:keychain:openai-key], and the real value is injected at the tool-execution edge, after the model asks for it by reference. It's more airtight — no secrets in the context window at all — but it costs complexity at every tool boundary, and it breaks the model's flow in exactly the places where having the raw value makes it smarter.

Right now I keep the secret in live memory and scrub the copies. What's your take — scrub everything, or keep the agent's working memory intact?

Ankita is open source (MIT) at https://github.com/akyourowngames/A.N.K.I.T.A — the detector lives in src/security/secret-scrubber.mjs and the echo-protection side is desktop/electron/secret-history.mjs. Try to fool it. If you find a leak, that's the most useful comment you could leave.

Top comments (0)