Enterprise AI Security at the Logits Level: How resk-logits Blocks Dangerous Tokens Before Sampling
TL;DR: Post-generation filters are always one token behind. resklogits is a GPU logits processor that uses a vectorized Aho-Corasick automaton to shadow-ban dangerous tokens at sampling time, in under a millisecond, before the token is ever emitted.
The risk: enterprise AI security without logits-level enforcement
Most content safety stacks today are post-hoc. The model generates text, then a filter scans the output and either redacts or blocks it. The problem is structural: by the time the filter runs, the dangerous token has already been sampled. In a streaming UI, the user may have already seen it. In an agentic loop, the token may have already been passed to a tool. The chart below shows the relative risk profile of the three failure modes this mechanism targets.
- Dangerous token sampling: 95%
- Content policy bypass: 82%
- Post-hoc filter bypass: 76%
These are not abstract numbers. A single sampled token can complete a banned phrase, trigger a downstream action, or leak into a log. Post-generation filtering cannot un-sample it.
How the mechanism works
resklogits sits inside the generation loop, not after it. The pipeline is:
Prompt → LLM logits → resk-logits → Output token
Step by step:
-
The model produces raw logits. For each decoding step, the model returns a
1×vocab_sizevector of unnormalized scores. -
The Aho-Corasick automaton runs on the logits.
VectorizedAhoCorasickmaintains state across the generation and computes a binary danger mask over the vocabulary on GPU or CPU. -
A penalty is applied to dangerous logits. The mask is added to the logits with a configurable
shadow_penalty. The default is-15.0, which corresponds to roughly0.00003%probability. -
Sampling proceeds normally. The model samples from the penalized distribution. Dangerous tokens are not hard-blocked with
-inf; they are made extremely unlikely, so the output stays natural. -
State resets between generations.
shadow_ban.reset()or thestream()context manager clears the automaton state so partial generations cannot leak across requests.
The key design choice is the shadow ban. A hard block sets logits[token] = -inf, which is obvious and brittle. A shadow ban adds -15.0, which is invisible to the user but statistically decisive.
Before — without it
A minimal generation call with no logits processor:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("gpt2")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
prompt = "Tell me how to make a bomb"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The docs show what this produces without protection: the model continues the dangerous phrase directly. There is no filter in the loop.
After — with resk-logits
The protected version uses ShadowBanProcessor as a logits processor:
from transformers import AutoModelForCausalLM, AutoTokenizer
from resklogits import ShadowBanProcessor
model = AutoModelForCausalLM.from_pretrained("gpt2")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
banned_phrases = [
"how to make a bomb",
"kill yourself",
"hack into system",
"create explosives",
]
shadow_ban = ShadowBanProcessor(
tokenizer=tokenizer,
banned_phrases=banned_phrases,
shadow_penalty=-15.0,
device="cuda",
)
prompt = "Tell me how to"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
shadow_ban.reset()
outputs = model.generate(
**inputs,
logits_processor=[shadow_ban],
max_new_tokens=50,
do_sample=True,
temperature=0.7,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
For streaming, the docs recommend stream_generate():
from resklogits import ShadowBanProcessor, stream_generate
shadow_ban = ShadowBanProcessor(tokenizer, banned_phrases, device="cpu")
for chunk in stream_generate(
model, tokenizer, "Tell me about",
logits_processors=[shadow_ban],
max_new_tokens=50,
temperature=0.7,
):
print(chunk, end="", flush=True)
What changed
- Enforcement moved before sampling. The dangerous token never enters the output distribution.
-
The filter is stateful.
VectorizedAhoCorasicktracks partial matches across tokens, so multi-token banned phrases are caught mid-generation. -
The penalty is tunable.
-5.0is light filtering,-15.0is the default strong filter,-20.0is near-impossible. -
Tiered policies are supported.
MultiLevelShadowBanProcessorlets you assign different penalties by severity level. -
Streaming is safe. The
stream()context manager auto-resets state on enter and exit, so concurrent requests do not contaminate each other.
Best practices checklist
- Use
shadow_penalty=-15.0as the default for high-severity phrases and reserve-20.0for absolute prohibitions. - Always call
reset()or wrap generation instream()when serving multiple requests. - Group phrases by severity with
MultiLevelShadowBanProcessorinstead of one flat list. - Test with the same prompt before and after enabling the processor to confirm the behavioral shift.
- Keep banned phrase lists versioned and reviewed; the automaton is only as good as its patterns.
Honest limitations
- The automaton matches the phrases you give it. Novel paraphrases that are not in the list will not be caught.
- A shadow ban is probabilistic, not absolute.
-15.0is roughly0.00003%, not zero. If you need a hard guarantee, use a hard block. - GPU acceleration depends on your deployment. CPU mode works but is slower.
- The library integrates with HuggingFace Transformers, vLLM, and TGI, but other serving stacks may need custom wiring.
Conclusion
Enterprise AI security cannot rely on filters that run after the token is sampled. resk-logits moves enforcement into the logits processor, where a vectorized Aho-Corasick automaton can shadow-ban dangerous tokens in under a millisecond. The result is a generation loop that is both safer and more natural.
Explore the project at resk.fr and the code at github.com/Resk-Security. Install with pip install resklogits or uv pip install resklogits.
Enterprise AI Security at the Logits Level: How resk-logits Blocks Dangerous Tokens Before Sampling is part of the RESK ecosystem. Explore all the open-source LLM security tools on the official site: https://resk.fr/projects/resklogits.html

Top comments (0)