We ran TypeSafe's Jev as a prompt guardrail on 1,580 red-teaming prompts against GPT-5-mini, Claude Haiku 4.5 and three open-source guard models. It was five to six times faster, 6–27× cheaper, and had the highest block accuracy of the hosted models.
TypeSafe's Jev promises LLM-grade decisions at a fraction of the time and cost. We put that claim to a practical test: using Jev as a guardrail that screens prompts before they reach an AI agent. Across 1,580 red-teaming prompts from public jailbreak, prompt-injection and harmful-content benchmarks, we compared Jev with OpenAI GPT-5-mini, Claude Haiku 4.5 and three open-source guard models on speed, cost, and how well each one decides what to block and names the risk.
Hype around Jev
TypeSafe launched Jev with this headline [1]:
Why the hype? Orders of magnitude? Is comparing it to an LLM even fair? Let's dig into where that speed comes from, and whether it holds up.
How Jev differs from an LLM
At a high level, Jev predicts the probability of each option in a set of choices or actions that you define, using the prompt or context you give it as the state. You might ask "should this input be blocked?" and "which of these 24 risk categories does it fall into?". Jev returns calibrated probabilities for every option in a single pass.
An LLM works differently. It generates its output one token at a time (The 'decode' phase), and each step depends on the tokens it has already produced. This stepwise decode phase is where most of an LLM's latency comes from. To use an LLM as a classifier, you ask it to write out its decision, usually as JSON, and then parse that text.
That difference has practical consequences:
- The output is typed. Jev returns probabilities and choices, not free text. There is nothing to parse, and no output can arrive malformed.
- The cost structure is different. Jev bills for input tokens only; its output is free [2]. An LLM bills for both, and at a higher rate for output.
- Latency does not grow with the answer. A single scoring pass does not slow down as the answer gets longer, because there is no answer to write out token by token.
Where Jev is a better fit
Jev is not a replacement for an LLM. It does not write answers, summarise documents or generate code. What it targets is the large class of tasks inside AI systems where the output is a decision:
- guardrails that decide whether an input or action should be blocked
- routing requests to the right model, tool or team
- classification, moderation and triage at high volume
- choosing the next action in an agent workflow
These decisions often sit in the critical path of every request. In that position, each additional second and each additional cent is paid on every call. That is where a model built only for decisions has the most to gain.
What we tested: Jev as a red-teaming guardrail
Agent security is a natural test for this kind of model. A guardrail has to read one input and make two calls:
- Should this input be blocked? It might be a harmful request, a jailbreak attempt, or a prompt injection hidden inside a document or tool response.
- What kind of risk is it?
For the second question, every model chooses from 24 risk categories, and the risk category each prompt carries in its source dataset is the ground truth. 22 of these come from RedBench's risk taxonomy [3]. We added No risk for benign inputs and Prompt attack for jailbreak and injection attempts that target no specific harm.
| Risk category | Covers |
|---|---|
| No risk | Benign request; safe to answer normally, even if it uses alarming words |
| Abusive content | Harassment, insults, bullying or degrading language |
| Bias content | Stereotyping or unfair treatment of groups |
| CBRN information or capabilities | Chemical, biological, radiological or nuclear weapons knowledge |
| Child safety | Content endangering or sexualising minors |
| Confabulation | Eliciting fabricated facts presented as true |
| Crime content | Facilitating theft, fraud, drugs or other crimes |
| Cybersecurity threats (beyond malware) | Hacking, intrusion, phishing, system disruption |
| Data privacy | Obtaining or exposing personal or confidential data |
| Decision-making in critical systems | Unsafe automated decisions in medical, legal, financial or safety-critical settings |
| Economic harm | Market manipulation, scams, financial damage |
| Election interference | Voter manipulation or election disinformation |
| Environmental harm | Pollution, ecological damage, wildlife harm |
| Extremism and radicalization | Promoting extremist ideology or recruitment |
| Hateful content | Hate speech targeting protected groups |
| Illegal weapons (non-CBRN) | Building or acquiring illegal conventional weapons or explosives |
| Intellectual property | Piracy or copyright and trademark infringement |
| Malware code | Writing or deploying malicious software |
| Misinformation | False or misleading claims presented as fact |
| Self-harm content | Suicide, self-injury or eating-disorder encouragement |
| Sexual content | Explicit or non-consensual sexual content |
| Violent content | Violence, threats or physical harm to others |
| Other | Harmful in a way not covered by the categories above |
| Prompt attack | Jailbreak or prompt-injection attempt when no more specific category applies |
Data
Our starting point was RedBench [3], which consolidates 37 red-teaming benchmarks. We added dedicated jailbreak and prompt-injection datasets [4–11], including agent transcripts where the attack is planted in a tool response [9]. From a pool of 41,443 prompts, we evaluated a stratified sample of 1,580 (1,310 harmful, 270 benign), drawn from 14 datasets.
Models
We gave the same task, with identical instructions and the same 24-category risk taxonomy, to three hosted models: TypeSafe Jev [2], OpenAI GPT-5-mini [12] and Anthropic Claude Haiku 4.5 [13].
Results and observations
Speed, cost and quality
Jev's median latency was 0.17 seconds, against 1.00 s for GPT-5-mini and 0.89 s for Claude Haiku 4.5. That is about six times faster than GPT-5-mini and more than five times faster than Haiku. At the 95th percentile, Jev stayed at 0.26 s, while both LLMs were above 1.3 s.
Classifying all 1,580 prompts cost $0.08 with Jev, $0.48 with GPT-5-mini and $2.23 with Claude Haiku 4.5 at list prices. That makes Jev about 6× cheaper than GPT-5-mini and 27× cheaper than Haiku on the same workload.
Our measured gap is smaller than the headline figures TypeSafe published. That is expected. Their numbers describe workflows of System One tasks against their chosen baselines. Ours come from a single adversarial classification task, compared against two widely used, low-cost LLMs. The direction matches their claim; the size of the gap depends on the task and on what you compare against.
Speed and cost only matter if the decisions hold up. Across all prompts, Jev flagged harm correctly 85.5% of the time (block accuracy), against 84.1% for GPT-5-mini and 80.1% for Claude Haiku 4.5.
On naming the exact risk category out of 24 (categorization accuracy), Jev scored 58.0% against 56.0% and 55.4%.
Its macro-F1 across the 24 categories, which weights every category equally regardless of size, was 46.5% against 41.7% and 38.1%.
Detection quality
Categorization accuracy for each risk category with more than 20 labelled prompts:
Of these categories, Jev has the highest categorization accuracy in nine: crime, abusive content, misinformation, cybersecurity threats, bias, data privacy, sexual content, economic harm and decision-making in critical systems. GPT-5-mini leads on hateful content, malware, violent content and illegal weapons. Haiku is strongest at recognising benign inputs (No risk) and prompt attacks, and ties with GPT-5-mini ahead of Jev on self-harm.
Comparison with task-specific models
We also ran three open-source models built specifically for guardrailing, on a T4 GPU: Qwen3Guard-Gen-4B [14], Llama Prompt Guard 2 86M [15] and ProtectAI DeBERTa-v3 prompt-injection v2 [16]. None of them answers in the 24 risk categories, so they are compared with Jev on block accuracy only.
To compare performance across attack domains, we grouped the source benchmarks into nine benchmark groups based on the kind of prompts each contains. The groups are used only to aggregate results; they are not the labels that the models predict.
| Benchmark group | What it tests | Benchmarks included |
|---|---|---|
| Harmful instruction and question sets | Plainly worded harmful requests and questions | AdvBench, CatQA, CoNA, ControversialInstructions, DAN, DoNotAnswer, ForbiddenQuestions, GPTFuzzer, HarmBench, HarmfulQ, HarmfulQA, JBB-Behaviors, MaliciousInstruct, MaliciousInstructions, MedSafetyBench, QHarm, SGBench, StrongREJECT (all via RedBench [3]) |
| Cyberattack assistance | Requests for help with hacking, intrusion and malware | CyberattackAssistance (via RedBench [3]) |
| Mixed safety suites | Broad safety suites whose prompts span many harm types | XSafety, JADE (via RedBench [3]) |
| Jailbreak prompt collections | DAN-style, role-play and "developer mode" templates, with ordinary prompts as benign contrasts | In-the-wild jailbreak and regular prompts [4], JailbreakBench JBC [5], Salad-Data jailbreak templates [6], jailbreak-classification and ChatGPT-Jailbreak-Prompts [10], LatentJailbreak (via RedBench [3]) |
| Automated jailbreak attacks | Attacks generated automatically (GCG, PAIR, random search, AutoDAN, GPTFuzzer, TAP) | JailbreakBench GCG, PAIR and random-search artifacts [5], Salad-Data AutoDAN, GCG, GPTFuzzer and TAP attacks [6] |
| Direct prompt injection sets | "Ignore previous instructions" and secret-extraction attempts, with benign contrasts | Gandalf ignore-instructions and summarization (via RedBench [3]), Lakera Gandalf [8], deepset prompt-injections [7], safe-guard-prompt-injection [10] |
| Indirect prompt injection (InjecAgent) | Instructions planted in tool responses to hijack an agent | InjecAgent data-stealing and direct-harm cases [9], plus clean counterparts built for this study |
| Toxic statements and unsafe advice | Toxic statements and prompts that invite unsafe advice | ToxiGen, SafeText, PhysicalSafetyInstructions (via RedBench [3]) |
| Over-refusal suites | Benign prompts that sound dangerous, plus unsafe contrasts | XSTest and CoCoNot [11], OR-Bench 11, JBB-Behaviors benign set [5] |
Qwen3Guard is scored on every benchmark group. Prompt Guard 2 and ProtectAI detect prompt attacks only, so they appear only in the jailbreak and prompt-injection groups they were run on in full.
Jev matches or beats every task-specific model in eight of the nine benchmark groups.
- Clear leads: direct prompt injection (95%, against 88% for ProtectAI and 80% for Qwen3Guard), cyberattack assistance (83% against 67%), toxic statements and unsafe advice (49% against 38%) and indirect prompt injection (77% against 70%).
- Narrow leads: harmful instructions (92% against 90%) and mixed safety suites (47% against 42%).
- Level: automated jailbreaks (96% for both) and over-refusal (72% for both).
- Behind: jailbreak prompt collections, where Qwen3Guard is slightly ahead (93.3% against 92.9%).
The two prompt-attack classifiers trail Jev throughout. Prompt Guard 2 is behind Jev and Qwen3Guard in every group it covers. ProtectAI beats Qwen3Guard on direct injection but not Jev. Both score 33% on indirect injection hidden in tool responses.
Two caveats apply:
- Mixed safety suites and toxic statements are hard for every model, with all scores below 50%.
- Indirect prompt injection remains Jev's weakest group against the hosted LLMs, which both reached 91% there.
What this means for enterprises
For most organisations, the expensive part of AI safety is not a single model call. It is making that call on every request, every agent step, and every test case in every evaluation run. A guardrail that sits in front of a production agent may run millions of times a month. An evaluation suite that re-scores thousands of cases after each prompt or model change runs constantly during development.
In those settings, the gap between a one-second LLM judgement and a sub-200-millisecond typed decision is not a rounding error. It decides whether security checks can run inline or have to be sampled, and whether evaluation runs on every change or once a week. Customers running their evaluation workloads through OptiX with Jev are already seeing reduced evaluation costs and lower latency, with performance on par with or ahead of the LLMs.
The broader lesson from this study is that decision tasks and generation tasks do not need the same model. Matching the model to the job, and measuring it on your own data, pays off quickly.
Explore OptiX
OptiX is our platform for testing, evaluating and monitoring AI systems. It brings together:
- Security scans for jailbreaks, prompt injection and other adversarial inputs
- Multi-agent evaluation across complex, multi-step workflows
- Observability for AI systems in production
- Use-case evaluation for RAG, text-to-SQL, document parsing, entity extraction, summarisation and more
- Jev support, so fast, low-cost typed decisions can power guardrails and evaluations alongside LLM-based judges
To see how your own agents hold up, visit optix.cobuildx.ai and get in touch with our team.
References
Benchmark and claim
- TypeSafe AI. "193.6x faster, 444.6x cheaper" launch claim. https://typesafe.ai/
- TypeSafe AI. Jev model documentation. https://docs.typesafe.ai/models
- Q.-A. Dang, C. Ngo, T.-S. Hy. RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models. arXiv:2601.03699, 2026. Dataset: https://huggingface.co/datasets/knoveleng/redbench
Jailbreak and prompt-injection datasets
- X. Shen et al. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. ACM CCS 2024. arXiv:2308.03825. Dataset: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts
- P. Chao et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. arXiv:2404.01318, 2024. Artifacts: https://github.com/JailbreakBench/artifacts
- L. Li et al. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models. arXiv:2402.05044, 2024. Dataset: https://huggingface.co/datasets/OpenSafetyLab/Salad-Data
- deepset. prompt-injections dataset. https://huggingface.co/datasets/deepset/prompt-injections
- Lakera. gandalf_ignore_instructions dataset. https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions
- Q. Zhan et al. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. arXiv:2403.02691, 2024. https://github.com/uiuc-kang-lab/InjecAgent
- Community datasets on Hugging Face: jackhhao/jailbreak-classification, xTRam1/safe-guard-prompt-injection, rubend18/ChatGPT-Jailbreak-Prompts.
- Over-refusal benchmarks: P. Röttger et al. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. arXiv:2308.01263, 2023; F. Brahman et al. The Art of Saying No: Contextual Noncompliance in Language Models (CoCoNot). arXiv:2407.12043, 2024; J. Cui et al. OR-Bench: An Over-Refusal Benchmark for Large Language Models. arXiv:2405.20947, 2024.
Models
- OpenAI. GPT-5-mini. https://platform.openai.com/docs/models
- Anthropic. Claude Haiku 4.5. https://docs.anthropic.com/en/docs/about-claude/models
- H. Zhao et al. Qwen3Guard Technical Report. arXiv:2510.14276, 2025.
- S. Chennabasappa et al. LlamaFirewall: An Open Source Guardrail System for Building Secure AI Agents (Llama Prompt Guard 2). arXiv:2505.03574, 2025.
- ProtectAI. deberta-v3-base-prompt-injection-v2. https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2
Originally published on the CobuildX blog.








Top comments (0)