DEV Community

Naveen Vikram
Naveen Vikram

Posted on

Your AI Guardrails Are a Taxonomy Pretending to Be an Ontology

A guardrail that only knows categories can label a problem. It can't reason about one.


The bouncer who never read the guest list

Most guardrail tools look the same: categories (content safety, PII, prompt injection), a detector for each, and a threshold. Cross it and you're blocked.

That's a taxonomy. It answers "which bucket does this belong in?" You need one, but a lot of teams (me included, a while back) treat the buckets as the whole safety system.

An ontology answers "what is this connected to, and what follows from that?" Think of a UPI payment. A taxonomy says "this is a payment." The system only approves it because it also knows who's paying, from which account, to whom, within what limit, and whether the PIN checks out.

Example 1: MedScribe's correction engine

Every medical term MedScribe touches ends up as corrected, review_required, flagged (two reasons) or ignored. On its own that's a taxonomy. But the label isn't read off the term. It's worked out from relationships:

  • If the raw term already matches a canonical name or alias, a short-circuit fires first: no correction needed.
  • Otherwise, retrieval similarity (≥ 0.90) and extraction confidence (≥ 0.70) must both clear.
  • Even then, if the top match and the runner-up are within 0.05, it goes to review. The engine is torn between two drugs, and it knows it.

(Honest note: those thresholds are engineering defaults, not validated on real clinical outcomes.)

Example 2: permissions that know what they're touching

A flat list says refund: allowed for admins. A typed rule says a refund belongs to a Payment, and is only allowed when that payment is in the right state and the caller has the right role. (That one's my illustration, not code from a project.)

Ken Huang makes this argument for agents: tools become verbs attached to types, and authorization depends on object state and relationships, not a flat function list. His bigger claim is that agent security will increasingly be judged at the ontology boundary, not the model boundary. You can't enforce a rule you can't represent.

Guardrails in the wild: Jev

TypeSafe AI's Jev is a good test case. It's a "System One" model: it can't write text, only return typed values with probabilities, from a fixed set of choices. Fast and cheap, which suits a guardrail.

The community project jev-guard uses it to vet an agent's tool calls before they run. Verdicts are allow, block or hold for a human, and if confidence drops below a threshold, it holds. That's the same abstain-when-unsure pattern as MedScribe.

But allow/block/hold are still just labels. What decides whether they're any good is the structured state you hand Jev: what the tool does, whether it's reversible, what the agent said it was trying to do. The model supplies the verdict; the ontology supplies the reasons. And its README is upfront that it's "not a security boundary," with about 68% accuracy on its task, so it can't replace sandboxing or least-privilege credentials.

Quick self-test

  1. When your guardrail blocks something, can it say which relationship made it unsafe, or only which category matched?
  2. Does the same action get a different answer depending on object state?

If your answers are "category" and "no," you've built a polite bouncer who checks IDs for the word "VIP" and has never seen the guest list.

Taxonomies aren't bad. They just shouldn't be the last step.


Sources: Ken Huang, "Why Ontology Matters for Agentic AI in 2026"; TypeSafe AI, "Introducing System One Models & Jev"; jev-guard (GitHub); BenchLM, "What Is Jev?".

Top comments (0)