Every day, developers, lawyers, and founders paste confidential source code, customer records, and internal business metrics directly into ChatGPT, Claude, and Gemini.
When you paste raw client data like "Sarah Jenkins from TechNova GmbH paid $45,000" into public LLM chats, that private information is sent across the internet, logged into provider telemetry databases, and potentially used to retrain future AI models.
In this guide, we will build a 100% client-side, zero-knowledge AI prompt sanitizer in pure JavaScript. It replaces confidential names, emails, and financial figures with structured placeholder tokens, and features a 1-click two-way restoration engine that maps real names back into AI responses without uploading a single byte to any external server.
1. Why Server-Side Redaction Destroys Privacy
Most online PII (Personally Identifiable Information) scrubbers operate by uploading your prompt to their own cloud servers to run machine learning models.
This creates a serious privacy trap: to protect your data from OpenAI, you are sending your confidential text to an unknown third-party server.
The Traditional Cloud Trap:
Your Prompt -> Third-Party Redaction Server -> OpenAI LLM
[--- Risk of Leaks / Logs at Every Step ---]
The Zero-Knowledge Client-Side Approach:
Your Prompt -> Local Browser RAM (Regex + Bracket Tokenization) -> OpenAI LLM
[--- 0 Network Requests / 0 Logs / 100% Private ---]
2. The Two-Way Masking Workflow
The secret to safe AI prompting is pseudonymization with reverse mapping:
-
Step 1 (Tokenize): Replace sensitive entities with typed tokens (
[PERSON_1],[COMPANY_1],[CURRENCY_1]). - Step 2 (Prompt AI): Send the sanitized text to ChatGPT or Claude. The AI understands the context perfectly and returns a response containing the tokens.
- Step 3 (Restore): Paste the AI reply back into your local browser tool. The in-memory dictionary swaps the tokens back into original names with one click.
Common Sensitive Data Types & Replacement Patterns
| Data Type | Example Real Value | Safe Token Replacement | Detection Method |
|---|---|---|---|
| Custom Secret Words | [Project Titan] |
[CUSTOM_1] |
Bracket [ ] Delimiter |
| Client Names | Sarah Jenkins |
[PERSON_1] |
Title Case Capitalization |
| Companies | Acme Corp / TechNova GmbH |
[COMPANY_1] |
Corporate Suffix Regex |
| Emails | sarah@technova.com |
[EMAIL_1] |
Standard RFC 5322 Regex |
| Phone Numbers | +1 (555) 234-5678 |
[PHONE_1] |
International Phone Pattern |
| Financial Amounts | $45,000 / €12,500 |
[CURRENCY_1] |
Currency Symbol Regex |
3. Implementing the JavaScript Anonymizer Engine
Here is the complete, self-contained JavaScript implementation that handles tokenization, state management, and reverse restoration directly in memory:
class PromptAnonymizer {
constructor() {
this.tokenMap = new Map(); // Maps [TOKEN] -> Real Value
this.counters = {
CUSTOM: 0,
EMAIL: 0,
PHONE: 0,
CURRENCY: 0,
COMPANY: 0,
PERSON: 0
};
}
// Generate unique token and store in local dictionary
createToken(type, originalValue) {
// Return existing token if value was already masked
for (let [token, val] of this.tokenMap.entries()) {
if (val.toLowerCase() === originalValue.toLowerCase()) {
return token;
}
}
this.counters[type] = (this.counters[type] || 0) + 1;
const token = `[${type}_${this.counters[type]}]`;
this.tokenMap.set(token, originalValue);
return token;
}
// Anonymize raw prompt text
anonymize(text) {
let output = text;
// 1. Manual Bracket Masking: [Any Secret Phrase]
output = output.replace(/\[([^\]]+)\]/g, (match, content) => {
return this.createToken('CUSTOM', content);
});
// 2. Email Addresses
const emailRegex = /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/g;
output = output.replace(emailRegex, (match) => {
return this.createToken('EMAIL', match);
});
// 3. Currency Figures ($45,000 / €100M / £500)
const currencyRegex = /([$€£¥₹])\s?(\d{1,3}(?:,\d{3})*(?:\.\d+)?(?:k|m|b|K|M|B)?)/g;
output = output.replace(currencyRegex, (match) => {
return this.createToken('CURRENCY', match);
});
// 4. Corporate Entities (Inc, LLC, GmbH, Ltd)
const companyRegex = /\b([A-Z][A-Za-z0-9&.-]+(?:\s+[A-Z][A-Za-z0-9&.-]+)*\s+(?:GmbH|LLC|Inc|Corp|Ltd|AG|SA|Pvt))\b/g;
output = output.replace(companyRegex, (match) => {
return this.createToken('COMPANY', match);
});
return output;
}
// Restore AI response back to original identities
restore(aiResponseText) {
let restored = aiResponseText;
for (let [token, realValue] of this.tokenMap.entries()) {
// Replace all instances of [TOKEN] globally
restored = restored.split(token).join(realValue);
}
return restored;
}
}
4. Testing the Engine: Step-by-Step Execution
Let's test our engine on a real confidential business prompt:
const sanitizer = new PromptAnonymizer();
const confidentialPrompt =
"Write an email to [Sarah Jenkins] at TechNova GmbH thanking her for the $45,000 project payment. Contact her at sarah@technova.com.";
// 1. Sanitize before sending to AI
const safePrompt = sanitizer.anonymize(confidentialPrompt);
console.log("Safe Prompt for AI:\n", safePrompt);
// 2. Simulated response returned by ChatGPT / Claude
const mockAiResponse =
"Subject: Thank you for your payment\n\nDear [CUSTOM_1],\n\nThank you for confirming the [CURRENCY_1] payment on behalf of [COMPANY_1]. We will reach out to [EMAIL_1] with next steps.";
// 3. Restore original identities locally in 1 click
const finalRestoredText = sanitizer.restore(mockAiResponse);
console.log("\nFinal Restored Email:\n", finalRestoredText);
Execution Output:
Safe Prompt for AI:
Write an email to [CUSTOM_1] at [COMPANY_1] thanking her for the [CURRENCY_1] project payment. Contact her at [EMAIL_1].
Final Restored Email:
Subject: Thank you for your payment
Dear Sarah Jenkins,
Thank you for confirming the $45,000 payment on behalf of TechNova GmbH. We will reach out to sarah@technova.com with next steps.
5. Security & Architectural Benefits
- Deterministic Accuracy: Regex and bracket delimiter parsing runs in < 2 milliseconds directly in browser V8 engine RAM.
-
Grammar Preservation: By using typed semantic placeholders like
[COMPANY_1]and[PERSON_1], the LLM maintains perfect grammar and pronoun alignment (he/she/they). - Zero Telemetry Liability: Because no tokens leave the client device, your workflows remain fully compliant with GDPR, HIPAA, and SOC 2 pre-prompt guidelines.
Discussion & Questions
How does your team protect sensitive customer data and credentials when using generative AI tools? Do you use browser extensions, manual search-and-replace, or strict data loss prevention (DLP) rules? Let's discuss in the comments below!


Top comments (0)