A guardrail that only knows categories can label a problem. It can't reason about one.
The bouncer who never read the guest list
Open the docs of almost any guardrail tool and you'll see the same thing: a list of categories. Content safety, prompt injection, PII, compliance. Each one has a detector, each detector has a threshold, and if something crosses the threshold, it gets blocked.
That's a taxonomy. It answers one question: "which bucket does this belong in?" You do need one. But a lot of teams (including me, a while back) end up treating the buckets as the whole safety system.
An ontology answers a different question: "what is this connected to, and what follows from that?" Easy to say, harder to feel. Think of a UPI payment. A taxonomy says "this is a payment." The system only approves it because it also knows who's paying, from which account, to whom, within what limit, and whether the PIN checks out. So here are two examples.
Example 1: MedScribe's correction engine
In MedScribe, every medical term the pipeline touches ends up with one of five logged outcomes: corrected, review_required, flagged (for two different reasons), or ignored. If that were the whole story, it's just a taxonomy.
But the outcome isn't decided by looking at the term alone. It's worked out from how the term relates to other things:
- If the raw term already matches a canonical name or alias, a short-circuit rule fires first. No correction needed, log it as identity-protected and move on.
- Otherwise, retrieval similarity (≥ 0.90) and extraction confidence (≥ 0.70) both have to clear before a correction is even allowed.
- Even then, if the top match and the runner-up are within 0.05 of each other, the correction goes to review. The engine is torn between two plausible drugs, and it knows it.
So the label is a conclusion, not a property of the term. I didn't call it an ontology when I built it, but that's the kind of thinking it is. (Honest note: those thresholds are engineering-judgment defaults. I haven't validated them on real clinical outcomes.)
Example 2: permissions that know what they're touching
A flat permission list says refund: allowed for the admin role. That's it.
A typed rule says something richer: a refund is an action that belongs to a Payment, and it's only allowed when that payment is in the right state and the caller has the right role. (This one is my illustration of the pattern, not code from a specific project.)
Ken Huang makes this exact argument in his 2026 piece on ontology and agentic AI: tools become verbs attached to types, and authorization depends on the state of the object and its relationships, not a flat list of functions. His bigger claim is that security and compliance for agents will increasingly be judged at the ontology boundary, not the model boundary.
That matches what I've seen building payment and clinical systems. You can't enforce a rule you can't represent. The first time an agent needs a constraint your schema has no word for, the guardrail has nothing to check.
"So do I have to model the whole enterprise?"
No. This is where most ontology advice goes wrong. Atlan's design guidance makes the sensible point that an ontology starting from one domain's real questions ships in weeks, while a whole-enterprise diagram rarely ships at all.
Pick the one decision your agent could get badly wrong, and model only what that decision depends on. MedScribe's version is a dictionary, an alias table, three numbers and one short-circuit rule. Small, but enough.
Quick self-test
- When your guardrail blocks something, can it say which relationship made the action unsafe, or only which category matched?
- Does the same action get a different answer depending on object state (a settled payment vs. a pending one)?
- When you add a new tool, do you extend the rules by adding a type, or by writing another keyword list?
If your answers are "category," "no," and "keyword list," you've built a polite bouncer who checks IDs for the word "VIP" and has never seen the guest list.
Taxonomies aren't bad. Classification is a fine first step. It just shouldn't be the last one.
Sources: Ken Huang, "Why Ontology Matters for Agentic AI in 2026" (Substack); OvalEdge, "Ontology vs Taxonomy in 2026"; Atlan, "Active Ontology Design for AI: A 2026 Enterprise Framework"; Maxim, "The Complete AI Guardrails Implementation Guide for 2026".

Top comments (0)