DEV Community

Emmanuel R for CobuildX AI

Posted on Originally published at cobuildx.ai

Jev as a Red-Teaming Guardrail: Putting TypeSafe's Model to Work in AI Security

We ran TypeSafe's Jev as a prompt guardrail on 1,580 red-teaming prompts against GPT-5-mini, Claude Haiku 4.5 and three open-source guard models. It was five to six times faster, 6–27× cheaper, and had the highest block accuracy of the hosted models.

TypeSafe's Jev promises LLM-grade decisions at a fraction of the time and cost. We put that claim to a practical test: using Jev as a guardrail that screens prompts before they reach an AI agent. Across 1,580 red-teaming prompts from public jailbreak, prompt-injection and harmful-content benchmarks, we compared Jev with OpenAI GPT-5-mini, Claude Haiku 4.5 and three open-source guard models on speed, cost, and how well each one decides what to block and names the risk.

Hype around Jev

TypeSafe launched Jev with this headline [1]:

TypeSafe's launch claim: 193.6x faster, 444.6x cheaper

Why the hype? Orders of magnitude? Is comparing it to an LLM even fair? Let's dig into where that speed comes from, and whether it holds up.

How Jev differs from an LLM

At a high level, Jev predicts the probability of each option in a set of choices or actions that you define, using the prompt or context you give it as the state. You might ask "should this input be blocked?" and "which of these 24 risk categories does it fall into?". Jev returns calibrated probabilities for every option in a single pass.

An LLM works differently. It generates its output one token at a time (The 'decode' phase), and each step depends on the tokens it has already produced. This stepwise decode phase is where most of an LLM's latency comes from. To use an LLM as a classifier, you ask it to write out its decision, usually as JSON, and then parse that text.

That difference has practical consequences:

  • The output is typed. Jev returns probabilities and choices, not free text. There is nothing to parse, and no output can arrive malformed.
  • The cost structure is different. Jev bills for input tokens only; its output is free [2]. An LLM bills for both, and at a higher rate for output.
  • Latency does not grow with the answer. A single scoring pass does not slow down as the answer gets longer, because there is no answer to write out token by token.

Where Jev is a better fit

Jev is not a replacement for an LLM. It does not write answers, summarise documents or generate code. What it targets is the large class of tasks inside AI systems where the output is a decision:

  • guardrails that decide whether an input or action should be blocked
  • routing requests to the right model, tool or team
  • classification, moderation and triage at high volume
  • choosing the next action in an agent workflow

These decisions often sit in the critical path of every request. In that position, each additional second and each additional cent is paid on every call. That is where a model built only for decisions has the most to gain.

What we tested: Jev as a red-teaming guardrail

Agent security is a natural test for this kind of model. A guardrail has to read one input and make two calls:

  1. Should this input be blocked? It might be a harmful request, a jailbreak attempt, or a prompt injection hidden inside a document or tool response.
  2. What kind of risk is it?

For the second question, every model chooses from 24 risk categories, and the risk category each prompt carries in its source dataset is the ground truth. 22 of these come from RedBench's risk taxonomy [3]. We added No risk for benign inputs and Prompt attack for jailbreak and injection attempts that target no specific harm.

Risk category Covers
No risk Benign request; safe to answer normally, even if it uses alarming words
Abusive content Harassment, insults, bullying or degrading language
Bias content Stereotyping or unfair treatment of groups
CBRN information or capabilities Chemical, biological, radiological or nuclear weapons knowledge
Child safety Content endangering or sexualising minors
Confabulation Eliciting fabricated facts presented as true
Crime content Facilitating theft, fraud, drugs or other crimes
Cybersecurity threats (beyond malware) Hacking, intrusion, phishing, system disruption
Data privacy Obtaining or exposing personal or confidential data
Decision-making in critical systems Unsafe automated decisions in medical, legal, financial or safety-critical settings
Economic harm Market manipulation, scams, financial damage
Election interference Voter manipulation or election disinformation
Environmental harm Pollution, ecological damage, wildlife harm
Extremism and radicalization Promoting extremist ideology or recruitment
Hateful content Hate speech targeting protected groups
Illegal weapons (non-CBRN) Building or acquiring illegal conventional weapons or explosives
Intellectual property Piracy or copyright and trademark infringement
Malware code Writing or deploying malicious software
Misinformation False or misleading claims presented as fact
Self-harm content Suicide, self-injury or eating-disorder encouragement
Sexual content Explicit or non-consensual sexual content
Violent content Violence, threats or physical harm to others
Other Harmful in a way not covered by the categories above
Prompt attack Jailbreak or prompt-injection attempt when no more specific category applies

Data

Our starting point was RedBench [3], which consolidates 37 red-teaming benchmarks. We added dedicated jailbreak and prompt-injection datasets [4–11], including agent transcripts where the attack is planted in a tool response [9]. From a pool of 41,443 prompts, we evaluated a stratified sample of 1,580 (1,310 harmful, 270 benign), drawn from 14 datasets.

Models

We gave the same task, with identical instructions and the same 24-category risk taxonomy, to three hosted models: TypeSafe Jev [2], OpenAI GPT-5-mini [12] and Anthropic Claude Haiku 4.5 [13].

Results and observations

Speed, cost and quality

Jev's median latency was 0.17 seconds, against 1.00 s for GPT-5-mini and 0.89 s for Claude Haiku 4.5. That is about six times faster than GPT-5-mini and more than five times faster than Haiku. At the 95th percentile, Jev stayed at 0.26 s, while both LLMs were above 1.3 s.

Latency per prompt

Classifying all 1,580 prompts cost $0.08 with Jev, $0.48 with GPT-5-mini and $2.23 with Claude Haiku 4.5 at list prices. That makes Jev about 6× cheaper than GPT-5-mini and 27× cheaper than Haiku on the same workload.

Cost per prompt

Our measured gap is smaller than the headline figures TypeSafe published. That is expected. Their numbers describe workflows of System One tasks against their chosen baselines. Ours come from a single adversarial classification task, compared against two widely used, low-cost LLMs. The direction matches their claim; the size of the gap depends on the task and on what you compare against.

Speed and cost only matter if the decisions hold up. Across all prompts, Jev flagged harm correctly 85.5% of the time (block accuracy), against 84.1% for GPT-5-mini and 80.1% for Claude Haiku 4.5.

Overall block accuracy

On naming the exact risk category out of 24 (categorization accuracy), Jev scored 58.0% against 56.0% and 55.4%.

Overall categorization accuracy

Its macro-F1 across the 24 categories, which weights every category equally regardless of size, was 46.5% against 41.7% and 38.1%.

Overall macro-F1

Detection quality

Categorization accuracy for each risk category with more than 20 labelled prompts:

Categorization accuracy by risk category, categories with more than 20 labelled prompts

Of these categories, Jev has the highest categorization accuracy in nine: crime, abusive content, misinformation, cybersecurity threats, bias, data privacy, sexual content, economic harm and decision-making in critical systems. GPT-5-mini leads on hateful content, malware, violent content and illegal weapons. Haiku is strongest at recognising benign inputs (No risk) and prompt attacks, and ties with GPT-5-mini ahead of Jev on self-harm.

Comparison with task-specific models

We also ran three open-source models built specifically for guardrailing, on a T4 GPU: Qwen3Guard-Gen-4B [14], Llama Prompt Guard 2 86M [15] and ProtectAI DeBERTa-v3 prompt-injection v2 [16]. None of them answers in the 24 risk categories, so they are compared with Jev on block accuracy only.

To compare performance across attack domains, we grouped the source benchmarks into nine benchmark groups based on the kind of prompts each contains. The groups are used only to aggregate results; they are not the labels that the models predict.

Benchmark group What it tests Benchmarks included
Harmful instruction and question sets Plainly worded harmful requests and questions AdvBench, CatQA, CoNA, ControversialInstructions, DAN, DoNotAnswer, ForbiddenQuestions, GPTFuzzer, HarmBench, HarmfulQ, HarmfulQA, JBB-Behaviors, MaliciousInstruct, MaliciousInstructions, MedSafetyBench, QHarm, SGBench, StrongREJECT (all via RedBench [3])
Cyberattack assistance Requests for help with hacking, intrusion and malware CyberattackAssistance (via RedBench [3])
Mixed safety suites Broad safety suites whose prompts span many harm types XSafety, JADE (via RedBench [3])
Jailbreak prompt collections DAN-style, role-play and "developer mode" templates, with ordinary prompts as benign contrasts In-the-wild jailbreak and regular prompts [4], JailbreakBench JBC [5], Salad-Data jailbreak templates [6], jailbreak-classification and ChatGPT-Jailbreak-Prompts [10], LatentJailbreak (via RedBench [3])
Automated jailbreak attacks Attacks generated automatically (GCG, PAIR, random search, AutoDAN, GPTFuzzer, TAP) JailbreakBench GCG, PAIR and random-search artifacts [5], Salad-Data AutoDAN, GCG, GPTFuzzer and TAP attacks [6]
Direct prompt injection sets "Ignore previous instructions" and secret-extraction attempts, with benign contrasts Gandalf ignore-instructions and summarization (via RedBench [3]), Lakera Gandalf [8], deepset prompt-injections [7], safe-guard-prompt-injection [10]
Indirect prompt injection (InjecAgent) Instructions planted in tool responses to hijack an agent InjecAgent data-stealing and direct-harm cases [9], plus clean counterparts built for this study
Toxic statements and unsafe advice Toxic statements and prompts that invite unsafe advice ToxiGen, SafeText, PhysicalSafetyInstructions (via RedBench [3])
Over-refusal suites Benign prompts that sound dangerous, plus unsafe contrasts XSTest and CoCoNot [11], OR-Bench 11, JBB-Behaviors benign set [5]

Qwen3Guard is scored on every benchmark group. Prompt Guard 2 and ProtectAI detect prompt attacks only, so they appear only in the jailbreak and prompt-injection groups they were run on in full.

Block accuracy by benchmark group, Jev and task-specific models

Jev matches or beats every task-specific model in eight of the nine benchmark groups.

  • Clear leads: direct prompt injection (95%, against 88% for ProtectAI and 80% for Qwen3Guard), cyberattack assistance (83% against 67%), toxic statements and unsafe advice (49% against 38%) and indirect prompt injection (77% against 70%).
  • Narrow leads: harmful instructions (92% against 90%) and mixed safety suites (47% against 42%).
  • Level: automated jailbreaks (96% for both) and over-refusal (72% for both).
  • Behind: jailbreak prompt collections, where Qwen3Guard is slightly ahead (93.3% against 92.9%).

The two prompt-attack classifiers trail Jev throughout. Prompt Guard 2 is behind Jev and Qwen3Guard in every group it covers. ProtectAI beats Qwen3Guard on direct injection but not Jev. Both score 33% on indirect injection hidden in tool responses.

Two caveats apply:

  • Mixed safety suites and toxic statements are hard for every model, with all scores below 50%.
  • Indirect prompt injection remains Jev's weakest group against the hosted LLMs, which both reached 91% there.

What this means for enterprises

For most organisations, the expensive part of AI safety is not a single model call. It is making that call on every request, every agent step, and every test case in every evaluation run. A guardrail that sits in front of a production agent may run millions of times a month. An evaluation suite that re-scores thousands of cases after each prompt or model change runs constantly during development.

In those settings, the gap between a one-second LLM judgement and a sub-200-millisecond typed decision is not a rounding error. It decides whether security checks can run inline or have to be sampled, and whether evaluation runs on every change or once a week. Customers running their evaluation workloads through OptiX with Jev are already seeing reduced evaluation costs and lower latency, with performance on par with or ahead of the LLMs.

The broader lesson from this study is that decision tasks and generation tasks do not need the same model. Matching the model to the job, and measuring it on your own data, pays off quickly.

Explore OptiX

OptiX is our platform for testing, evaluating and monitoring AI systems. It brings together:

  • Security scans for jailbreaks, prompt injection and other adversarial inputs
  • Multi-agent evaluation across complex, multi-step workflows
  • Observability for AI systems in production
  • Use-case evaluation for RAG, text-to-SQL, document parsing, entity extraction, summarisation and more
  • Jev support, so fast, low-cost typed decisions can power guardrails and evaluations alongside LLM-based judges

To see how your own agents hold up, visit optix.cobuildx.ai and get in touch with our team.

References

Benchmark and claim

Jailbreak and prompt-injection datasets

Models


Originally published on the CobuildX blog.

Top comments (0)