Your coding agent just read your passwords and API keys. Its next request goes to OpenAI or Anthropic. I built a local check so they don't go with it.
How I got here
Lately, every hardware pitch has "local AI" somewhere in it. I use AI heavily in my engineering workflow, so naturally I was intrigued. If "AI-ready" is going to be part of the sales pitch, I'd at least like it to do something beyond helping justify the price tag.
I tried small local models for tasks I normally give hosted frontier models. I kept having to supervise and correct them enough that I wasn't getting the benefit I wanted. Where would they actually fit?
Then I noticed a trade-off with hosted models: whatever I send them leaves my machine. When a coding agent reads a file, the file's contents become part of its next request to the model. An API key or password it didn't need goes along for the ride.
Could a small local model check that output first? Not a replacement for the hosted model, just a second pair of eyes.
How the check works
I built it as a plugin for Torana, an open-source, local-first proxy that sits between your coding agent and its model provider.
- Your agent (Claude Code, Codex, Gemini or another harness) reads a file or runs a tool on your machine.
- Before the next request leaves, Torana's
piiplugin sends the new tool output to a local model you choose. - If the model flags an API key, password or private personal data, the plugin replaces the whole tool result with a short explanation. The hosted model never sees the original.
- The agent can acknowledge that and try another approach, for example reading a narrower range.
The scanner only sees the tool name and the numbered output, not your whole conversation. Earlier decisions are replayed on later turns rather than rescanned.
It's an extra check, not complete protection. A scanner can miss things, and it can block harmless output.
Which small models were actually good at this?
I compared seven local chat models and several specialised detectors on 60 sensitive and 60 harmless synthetic examples. The chat models used Q8 GGUF weights on an Apple Silicon Mac with 24 GB of memory, through llama.cpp.
| Model / setup | Detected / 60 sensitive | Harmless blocked / 60 | Invalid / 120 |
|---|---|---|---|
| Qwen2.5 3B | 60 | 18 | 0 |
| Qwen3.5 0.8B | 60 | 42 | 0 |
| privacy-filter (OpenAI, reference runtime) | 58 | 9 | 0 |
| SecretMasker (DistilBERT, threshold 0.5) | 57 | 10 | 0 |
| Adapted Ettin17M (research pilot) | 57 | 18 | 0 |
| Laya | 55 | 28 | 0 |
| Ettin17M PII | 55 | 33 | 0 |
| Ettin68M PII | 54 | 37 | 0 |
| LFM2.5 350M | 48 | 60 | 24 |
| SecretMasker (threshold 0.99) | 43 | 3 | 0 |
| pii_guard (regex baseline) | 24 | 0 | 0 |
| Qwen2.5 0.5B | 5 | 60 | 110 |
| Falcon-H1-Tiny 90M | 0 | 3 | 3 |
| Qwen3 0.6B | 0 | 0 | 0 |
| LFM2.5 1.2B | 0 | 60 | 120 |
A few things surprised me:
- Blocking isn't the same as detecting. When a model returned garbage, the plugin blocked the output to be safe. Those count as harmless results blocked, never as detections. Qwen2.5 0.5B "blocked everything" mostly because 110 of its 120 answers were invalid.
- Valid JSON doesn't mean correct. Qwen3 0.6B produced perfectly structured answers and missed all 60 secrets.
- Size isn't everything. Qwen2.5 3B caught all 60, but so did Qwen3.5 0.8B, which also blocked far more harmless output.
- Prompt wording swings results both ways. On a second test set built from real open-source code with planted secrets, Qwen2.5 3B caught 50 of 52 with the full prompt. Two simpler yes/no prompts landed at 37 and 49, both without blocking any harmless code.
- Specialised detectors are fast but brittle. SecretMasker runs in about 11 ms and caught all 52 real-code secrets, but it also blocked 12 of 26 harmless files, often mistaking opaque hashes for keys.
There's no winner badge here. These are small, synthetic test sets, and later plugin versions need fresh evaluation. The full write-up covers the methodology and its limits.
Try it yourself
This uses Qwen2.5 3B Instruct Q8 as a starting point (about 3.3 GB to download). Check its license before using it at work.
1. Start the local scanner with llama.cpp (brew install llama.cpp on macOS or Linux, winget install llama.cpp on Windows):
llama-server -hf bartowski/Qwen2.5-3B-Instruct-GGUF:Q8_0 --alias local-pii \
--host 127.0.0.1 --port 8081 -c 8192 -np 1 --jinja --no-context-shift
2. Install Torana and the plugin (macOS or Linux; Windows instructions are in the quickstart):
curl -fsSL https://torana.sh/install.sh | sh
torana start --port 8143
torana plugin install pii
torana open
3. Connect the model to the plugin in the browser UI that opens:
-
Settings → Add provider: name
local-scanner, upstream URLhttp://127.0.0.1:8081/v1, formatopenai, No authentication, default modellocal-pii. Save. -
Pipeline → pii → Resource bindings and limits → scanner: choose
local-scanner. - Review the permissions and choose Approve & enable.
4. Route your agent through Torana and test with a fake key:
echo 'PAYMENT_API_KEY=sk_test_torana_demo_not_a_real_key_123' > demo-sensitive.txt
echo 'RETRY_COUNT=3' > demo-safe.txt
ANTHROPIC_BASE_URL="$(torana endpoint anthropic)" claude
Ask it to read demo-sensitive.txt. You should see Tool output withheld in the reply and in Torana's Live Feed. demo-safe.txt should come through normally. Codex, Gemini and other harnesses have setup recipes here.
Scanning adds about a second per tool result, and large outputs that don't fit the model's context are blocked as scan failures. Start with small files. When you're done, torana stop --yes.
Why a proxy, and not just a hook?
In typical engineer fashion, I overengineered the reverse proxy I needed for this. It became Torana: reusable plugins, support for different model API formats, and hooks into the agent's traffic. The useful part of that detour: I can build a plugin once and use it across Claude Code, Codex, Gemini and others, instead of rebuilding it for each one.
There are plugins for usage logging, tool policy and context compaction too. The plugin SDK is in Go and Rust, sandboxed in WASM.
Help make it better
If you try it, I'd love to hear what breaks:
- harmless code that gets blocked;
- secrets that get missed;
- a local model or detector that does better.
Open an issue with a synthetic example and your setup (please never real credentials). And if you've built a little harness-specific tool or workflow hack, try turning it into a Torana plugin.
- Site: torana.sh
- GitHub: torana-edge/torana-edge

Top comments (0)