Your AI endpoint is probably the most expensive route in your app. Every request to /api/chat turns into tokens you pay for, and every MCP tool call can read your database, export records or kick off work in another system.
Most apps protect these routes the same way they protect everything else: authentication, then a rate limit that counts requests. That works until someone signs up, gets a valid session, and points a script at your chat route. Or until an agent with a perfectly valid token calls your documents.export tool sixty times in a minute. Request counts treat a one-row search and a full export as the same thing.
We built AI Protection for that gap, and it is now in public beta.
What it is
AI Protection is a set of server-side SDKs you install in your own backend. They sit inside your route, after your authentication and input validation and immediately before the model or tool runs. The rule is simple: if a request is denied, the model or tool never starts.
- Node.js / Next.js, Go and Python (FastAPI) SDKs for request admission, plus an experimental Cloudflare Workers adapter
- A Node MCP adapter for Streamable HTTP tool servers
- A dashboard where you see decisions, model attempts, tools and individual callers, and can pause a caller or a tool
The SDKs are public and Apache-2.0: Node, Go, Python.
Protecting a chat route (Node)
Create the instance once in a server-only module:
import { createAIProtection } from '@webdecoy/ai-protection';
const protect = createAIProtection({
webdecoyUrl: process.env.WEBDECOY_URL!,
webdecoyKey: process.env.WEBDECOY_KEY!,
propertyId: process.env.WEBDECOY_PROPERTY_ID!,
subjectSecret: process.env.WEBDECOY_SUBJECT_SECRET!, // random, 32+ characters
scopeId: 'support-chat',
route: '/api/chat', // a fixed route template, never a raw URL
protectionMode: 'observe',
resolveClientIP: trustedClientIP, // your ingress decides which IP to trust
});
Then, inside your route, after auth and validation:
return protect(request, () => callYourModel(request.signal));
trustedClientIP and callYourModel are your code, not SDK exports. The callback returns a normal Response, and a streaming response passes through unchanged. Start in observe mode, look at what WebDecoy would have done, and switch to enforce when you trust it.
Protecting MCP tools
If your AI feature exposes tools over MCP, the adapter wraps each tool so every call is authenticated, validated and authorized by your code before it runs:
import { createProtectedMCPHandler } from '@webdecoy/ai-protection/mcp';
const handler = createProtectedMCPHandler({
resource: 'https://api.example.com/mcp',
authorizationServer: 'https://your-issuer.example.com/',
policyVersion: 'records_v1',
authenticate: (request, { signal }) => verifyAccessToken(request, { signal }),
tools: {
'records.read': {
description: 'Read an owned record',
inputSchema: {
type: 'object',
properties: { id: { type: 'string' } },
required: ['id'],
additionalProperties: false,
},
requiredScopes: ['records:read'],
validate: (args) =>
args !== null && typeof args === 'object' && !Array.isArray(args)
&& Object.keys(args).length === 1 && 'id' in args && typeof args.id === 'string',
authorize: ({ caller, args }) => records.canRead(caller, args),
execute: ({ caller, args, signal }) => records.readForTenant(caller.tenant, args, signal),
},
},
});
A few details we cared about:
- Tool listing is filtered by scope, and every call checks scopes again. Naming a hidden tool directly does not get around it.
- Tenant comes from your verified caller, never from tool arguments, so a forged
tenantargument is rejected before your code runs. - Every tool result carries an action ID, and the dashboard can open that exact call.
Limits measured in work, not requests
A request counter cannot tell a one-row search from a full export. AI Protection can limit work: each tool declares a conservative maximum, the SDK reserves it across all your replicas before the call, and settles what your code actually did afterwards. Model budgets work the same way, reserving tokens before the provider call and settling confirmed usage.
You also pick what happens when WebDecoy is unreachable, per control. Fail open keeps serving with degraded coverage. Fail closed keeps a hard limit and refuses with a 503. Your own authorization rules apply either way.
Checking it without paying for a model
The part we are happiest with is the install doctor. From a checkout of the Node SDK:
git clone https://github.com/WebDecoy/ai-protection && cd ai-protection && npm ci
node scripts/doctor-mcp.mjs detect /path/to/your/app
detect reads your project (it never runs your code), tells you whether the layout is supported, and prints the exact setup commands with your files filled in. Each command has a plan mode that changes nothing, and a rollback.
Then doctor-mcp.mjs check plan.json runs your own synthetic calls against the running route: what should be allowed, what should be forbidden, what should be refused as cross-tenant, and what happens on cancellation. Add an outageRuntime and it serves a stand-in WebDecoy that returns 503 to everything, so you can see your fail-open and fail-closed choices actually behave. For model routes, checkModelProtection does the same against a stub provider that counts its calls. No paid model calls anywhere.
What gets sent, and what does not
The SDKs send request metadata (method, your route template, client IP, User-Agent, header names), pseudonymous caller identifiers hashed with a secret you hold, decisions and outcomes, MCP tool names and schema fingerprints, and token counts for budgets.
They do not send prompts, model responses, tool arguments, tool results, raw user identities or credentials. The pseudonymous identifiers are correlatable on purpose, so we do not call them anonymous.
What it is not
This is a beta, so here is the honest list:
- It does not replace your authentication or authorization. It runs after them and relies on them.
- It is not prompt-injection or output-safety protection. It checks who is calling and how much work they are allowed, not what the prompt says.
- We have not published detection accuracy numbers. We will when there is real-world data behind them.
- The MCP adapter is Node only for now. Go and Python cover request admission.
- It requires a code change in your backend. There is no proxy you can drop in front.
Try it
- Product page: webdecoy.com/product/ai-protection
- Setup guide: docs.webdecoy.com/ai-protection/setup
- Install:
npm install @webdecoy/ai-protection,go get github.com/WebDecoy/ai-protection-go@v0.1.0-beta.0orpip install webdecoy-ai-protection==0.1.0b1
Start in observe mode on a test route, run the doctor, and tell us what breaks. We are especially interested in hearing from anyone running MCP servers in production.
Top comments (0)