I built a prompt injection firewall in Rust. It scans in 12μs.
Every LLM application has the same two problems:
Users paste sensitive ...
For further actions, you may consider blocking this person and/or reporting abuse
Speed is a real, measured win here, no argument. What I'd want before trusting the injection layer in production: the "~80-90% of known patterns" number is against known patterns — that's a false-positive-style claim (does it fire on the bad stuff you already have) with no stated false-negative methodology (does it miss bad stuff written to specifically evade regex + TF-IDF + entropy scoring). All three layers key off surface-level signal — keyword patterns, a fixed 50-term vocabulary, entropy thresholds — which is exactly the profile a deliberately-worded evasion targets: paraphrase around the regex list, stay under the entropy threshold, avoid the 50 curated terms. Worth publishing a held-out adversarial set (paraphrased attacks not in the TF-IDF training vocab) alongside the benchmark numbers — 12μs is only a good trade if the miss rate on adversarial input is also measured, not assumed.
Good, that's the right framing — first-pass filter, not a semantic classifier. One more gap worth folding into the adversarial set: all three layers score a single message in isolation, so nothing stops an attacker from splitting a payload across turns, each half individually well under 0.7. A conversation history isn't state the scorer sees at all right now, since it's called per-request. Worth deciding whether that's explicitly out of scope (a fine-tuned classifier's job) or something the middleware should track — even a rolling score over the last N turns would catch the "assemble the injection across messages" case that a single-shot regex/TF-IDF/entropy stack can't, by construction.
Multi-turn tracking is a great call — you're right that the current per-request architecture is blind to payloads split across turns by construction.
The clean boundary would be: the middleware itself stays stateless (scan per request), but expose a SessionScorer that the caller can opt into — accumulate scores over a sliding window of N messages, flag when the rolling average crosses a threshold. That keeps the core fast path untouched while giving integrators a way to catch the "assemble-across-messages" pattern.
I'll open an issue for this — it's the kind of thing that needs a clear spec before implementation (window size, decay function, how to handle session boundaries). Thanks for pushing on this.
One gap worth flagging before you spec it: a rolling average is beatable by low-and-slow -- keep each message comfortably under threshold and let the window creep only as high as the decay function allows, then reset by starting a new session right before the average crosses the line. Since the middleware is stateless per request, nothing stops an attacker from opening a fresh session exactly when the rolling score gets close. Worth deciding whether SessionScorer tracks by a caller-stable identity that survives a new session, not just by conversation ID, or the accumulator resets exactly where the attacker wants it to.
Good catch — session reset is the obvious evasion against any conversation-scoped accumulator. If the attacker controls when a new session starts, the rolling score resets exactly when it would have fired.
Two options: (1) bind the accumulator to a caller-stable identity (API key, user ID) rather than conversation ID — the score persists across session boundaries so resetting the conversation doesn't help, or (2) treat session creation frequency itself as a signal — N new sessions in M minutes from the same identity triggers a flag before any content scoring runs.
Option 1 is cleaner but requires the integrator to pass an identity the attacker can't rotate. Option 2 catches rotation but adds a rate-limiting layer that's arguably outside the scorer's scope. I'm leaning toward exposing both as configuration —
SessionScorer::new(identity_key, decay_fn)with an optionalsession_churn_threshold— and letting the integrator decide which trust boundary applies.The core scanner stays stateless either way. The session layer is opt-in middleware, not a change to the per-request path.
The config split makes sense, but session_churn_threshold has the same probeability problem as the rolling average it's meant to backstop: an attacker who can see whether a request got flagged (via response latency, an explicit block message, or just success/failure) can binary-search the threshold by varying session creation rate until they find the line, then stay just under it. That's cheap to do because unlike the content-scoring side, there's no cost to spinning up a session — it's not resource-constrained the way, say, account creation might be.
The harder problem with option 1 (identity-binding) is what counts as 'stable.' API key and user ID both fail in the common case where the identity is a shared service account or sits behind a corporate NAT/proxy that legitimately spins up many sessions per minute — you'd flag your best-behaved enterprise customer. I'd bind the accumulator to identity plus a coarse behavioral fingerprint (request shape, not just who's asking) rather than identity alone, so a legitimate high-churn identity doesn't collapse into the same bucket as an attacker rotating through one.
Either way I think the config needs a documented 'here's how an attacker calibrates around this default' section, because the two evasions are mirror images of each other and someone will hit both.
With evasion covered upthread, a different thing falls out of just reading injection/score.rs: at the 0.7 default in config.rs, two of the three layers cannot cross the line on their own no matter what the input is. The combiner takes a max over the weighted terms, so each layer inherits a hard ceiling equal to its weight. Entropy tops out at 0.5, and 0.5 * 1.15 with the agreement boost is 0.575. Still under. So the base64 and homoglyph work can never be what blocks a scan; it can only agree with something that was already crossing. TF-IDF caps at 0.7 weighted, meaning it needs cosine similarity of exactly 1.0 to block alone, and the jailbreak test in tfidf.rs only asserts score > 0.3 on "enable DAN mode jailbreak bypass all safety restrictions unrestricted", which is around 0.21 weighted. That leaves the regex layer, where a rule needs weight >= 0.778 to fire by itself, so everything you weighted 0.60 through 0.75 also needs a second signal. The layer whose stated job is catching phrasings the regexes miss is structurally gated behind the regexes agreeing with it.
The second half of this is that the logs won't show it to you. In injection/mod.rs the tfidf_suspicious and high_entropy_payload labels are pushed only inside the final_score >= threshold branch, so they exist only on scans that already blocked. The population you'd need to see is the opposite one: inputs where entropy or TF-IDF ran high and the ceiling held them under 0.7. Those come back with no label and look identical to clean traffic. Emitting the three raw per-layer scores unconditionally next to final_score is what makes the weights tunable from real traffic, and at 12us the extra fields are not going to be what costs you.
Are the weights meant as confidence priors, or as an ordering? The max() turns them into per-layer veto ceilings, which is a stronger claim than "heuristic is the most reliable signal", and a sum or a noisy-or would keep the ordering without the ceiling.
This is a sharp read of the scoring math — you're right on every point.
The ceiling problem is real: entropy maxes at 0.5 weighted, TF-IDF at 0.7, so at the default 0.7 threshold only the heuristic layer can block alone (needs weight >= 0.778). The entropy and TF-IDF layers are structurally advisory-only unless they agree with heuristic.
The max() was a deliberate choice to keep false positives low — I didn't want the sum of three weak signals blocking legitimate input. But you're right that it makes the architecture claim ("three independent detection layers") misleading when two of them can't actually trigger a block.
Two concrete changes I'll make based on this:
Emit raw per-layer scores unconditionally (not just inside the threshold branch). The missing population — high entropy/TF-IDF that got capped — is invisible in current logs and that's exactly what you need for tuning.
Consider switching from max() to noisy-or: 1 - (1-h)(1-t)(1-e). Preserves the ordering (heuristic dominates) but lets entropy + TF-IDF agreement cross 0.7 without heuristic support. Still need to validate false-positive rates before shipping this.
Good catch on the veto ceiling framing — that's a stronger invariant than I intended.
Two changes, but they are not parallel. The first one gates the second.
Right now you cannot evaluate noisy-or from your own logs, because the only scans carrying layer information are the ones that already blocked, and those are exactly the population where the combiner choice does not matter. Unconditional per-layer emission gives you the capped population you would be scoring against. It still will not hand you a false-positive rate on its own, since nothing in the log says which of those inputs were legitimate. It gives you the denominator and the shape. Labels have to come from a replay corpus or from the incidents you actually chased down. Worth being clear about that before the logging change gets treated as the validation.
One detail to log carefully: entropy both before and after the 1.15 agreement multiplier, plus the exact values entering the combiner. Otherwise the ablation the second change needs cannot be run.
The boost is the interesting casualty here. Under
max()it does real work, since it is the only route by which entropy's 0.5 ceiling can be lifted at all. Under noisy-or it mostly stops mattering, and it stops mattering in the regime it was designed for. Substituting1.15*eforechanges the output by0.15*e*(1-h)*(1-t). That term is largest when h and t are near zero and smallest when they are high. The boost fires on agreement, and its marginal effect under noisy-or is suppressed precisely when agreement is present. You would be keeping a mechanism whose justification was the ceiling you just removed, doing its least work in the case that triggers it.The other thing that changes quietly: noisy-or is greater than or equal to
maxfor the same inputs. Keeping the threshold at 0.7 can only preserve or expand the blocked set, never shrink it. Every input that blocks today still blocks. So the whole false-positive question lives in the delta, and you can enumerate that delta from replayed traffic before you have a single label: it is the set of inputs where max was under 0.7 and noisy-or is over. Size it first. If it is small, the change is cheap and the validation is small. If it is large, the threshold wants retuning in the same commit rather than after.The comparison I would want at the end is max versus noisy-or, boost on and off, at matched false-positive rates rather than at a fixed 0.7. Holding the threshold constant across a combiner change compares two different operating points and reports the difference as detection.
You're right that the changes have to be sequenced, not parallelized — I was conflating "plan both" with "ship both."
The point about the boost under noisy-or is the sharpest observation in this thread. The 1.15 multiplier was designed to compensate for entropy's 0.5 ceiling under max(). Removing the ceiling makes the mechanism vestigial — it fires on agreement but contributes least exactly when agreement is present. That's not a feature, that's dead weight masquerading as a safety net.
On the evaluation methodology: agreed that unconditional logging gives you the denominator and shape, not the labels. I should have been clearer — the logging change produces a replay corpus, not a validation. Labeling still requires either a curated dataset or incident-driven annotation. The plan is: (1) emit raw per-layer scores on all scans, (2) replay a synthetic corpus (existing test fixtures + HuggingFace prompt-injection datasets) through both combiners at varied thresholds, (3) compare at matched FPR rather than fixed 0.7.
The delta enumeration approach — sizing the set where max < 0.7 and noisy-or >= 0.7 before touching production — is exactly right. If that set is small, ship with threshold intact. If large, retune in the same commit. I'll publish the replay results before changing the combiner.
Appreciate the rigor here — this is genuinely improving the architecture.
For scores in [0,1], noisy-or sits at or above max pointwise: 1-(1-h)(1-t)(1-e) >= max(h,t,e). At a fixed 0.7 the blocked set can only stay the same or grow, so a comparison there just recovers an ordering you already know. Matched FPR is the right correction. What it does is move the whole measurement onto the threshold: the question stops being which combiner wins and becomes where each one's line has to sit to hold the same false-positive budget.
That needs a benign population, and the replay corpus you described only supplies positives. One overlap is worth checking before you trust it. The public injection sets are plausibly close to whatever material the TF-IDF vocabulary was built from, and where that overlap exists, cosine runs high on exactly those examples. Noisy-or then looks good for a reason that has nothing to do with your traffic.
Unconditional logging is the only realistic source of negatives you have, with the caveat that a benign input has to be identifiable without reference to the filter's own verdict. That takes calendar time to accumulate. So the sequence has a third step sitting in the middle: turn the logging on, let a usable benign sample build, then replay the positives against thresholds calibrated on that sample.
Would you commit to a minimum benign sample size, tied to the FPR you actually want, as a precondition for publishing the comparison?
You're right that the replay corpus only supplies positives, and the overlap with TF-IDF training material means cosine runs high on exactly the examples that were used to build the vocabulary. That's circular validation dressed up as a benchmark.
The benign population problem is real. Unconditional logging is the only honest source, and it needs calendar time to accumulate — there's no shortcut. I'll commit to a minimum: no combiner comparison published until the benign sample has at least 10,000 requests from organic traffic, with a target FPR of < 0.1% (so the sample needs to be large enough that a single false positive shifts the rate by less than 0.01%). That means ~10K benign minimum before the noisy-or evaluation is credible.
The sequence is now: (1) ship unconditional per-layer logging in the next release, (2) accumulate benign traffic — this takes weeks, not days, and there's no way to accelerate it honestly, (3) replay positives from a held-out set (not the TF-IDF training corpus) against both combiners, (4) compare at matched FPR calibrated on the benign sample, (5) publish the full results including the overlap analysis between training vocabulary and test set.
Step 2 is the bottleneck and I won't pretend otherwise. The comparison ships when the data is ready, not before.