DEV Community

Cover image for Grounding AI Agents in Business Meaning: Semantic Layers, Ontologies, and MCP Elicitation in Practice
Seenivasa Ramadurai
Seenivasa Ramadurai

Posted on

Grounding AI Agents in Business Meaning: Semantic Layers, Ontologies, and MCP Elicitation in Practice

When an AI agent gets it wrong, the model is usually not the problem. The agent was never taught what your terms mean. Semantic models and ontologies fix that, and here is a small working sketch.

We have all done it. You hand someone a task, they come back with something different from what you pictured, and the thought slips out: this person is dumb.

Then you think about what actually happened. They were not short on intelligence. They were short on context. They did not know which report you meant, which customers count as "active," or that "urgent" on your team means today, not this week. They did exactly what they understood. The gap was in shared meaning.

The same thing happens with AI agents. If you treat an agent as a digital worker, it fails the way a new hire fails, and it needs the same fix:** direction and teaching.**

The "dumb agent" problem

Take this task: "Send a follow up to all our top clients about the overdue invoices."

It reads as simple, but the agent has to settle questions you never said out loud.

  • Is a "client" an account, a contact, or a company?
  • What makes a client "top": revenue, contract size, or strategic value?
  • What is "overdue": one day past due, or past the grace period?
  • Which system is the source of truth: the CRM, the billing tool, or the spreadsheet finance keeps on the side?

A strong model will guess, and the guess will often look reasonable. Reasonable is not the same as correct. When it is wrong, we blame the agent, just as we would blame the new hire. The model was fine. The meaning was missing.

Intelligence is not the bottleneck

Current foundation models can reason, write, plan, and call tools. What they do not arrive with is your world: your business terms, your rules, how your concepts relate, and what is allowed.

A person absorbs that over months of onboarding, conversations, mistakes, and mentoring. An agent gets a prompt, a few documents, and a task, and then we are surprised when it misreads us.

So the useful question is not "how do I get a smarter agent?" *It is *"how do I teach this agent what things mean here?"

Semantic models: the vocabulary

A semantic model defines what your data and terms mean in business language. Instead of seeing a column called cust_stat_cd = 3, the agent sees "Active Customer" with a definition attached.

It is the glossary a good manager hands a new employee on day one.

Revenue means recognized revenue, not bookings.
An active customer has bought something in the last 90 days.
Region follows the sales territory map, not the shipping address.

With that in place the agent looks up what a term means instead of inventing an answer, the same way a well onboarded employee would.

Ontologies: how the pieces connect

A semantic model gives you the words. An ontology gives you the relationships between them the concepts in a domain, their properties, and how they link.

For the invoice example:

  • A Customer has one or more Accounts.
  • An Account receives Invoices.
  • An Invoice is overdue when today is later than its due date plus the grace period.
  • A Contact works for a Customer and can approve payments only if the role is Finance.

People pick this up slowly and rarely write it down. An ontology writes it down in a form a machine can use.

One thing to be clear about: an ontology on its own does not enforce anything. It is a description of the domain. Something has to read it and act on it, such as a knowledge graph query, a validation step in your workflow, or a tool the agent calls. The next section shows what that looks like.

What it looks like in code

This is a deliberately small sketch to show the shape, not a production design. Two pieces: a tool that serves definitions, and a rule check the agent calls before it acts.

First, the semantic model exposed as an MCP tool, so the agent looks terms up instead of guessing.

from typing import Literal

from mcp.server.fastmcp import Context, FastMCP
from pydantic import BaseModel, Field

mcp = FastMCP("business-glossary")

# Governed by the glossary owner. The agent can read it but never writes to it.
GLOSSARY = {
    "active customer": {
        "definition": "A customer with at least one purchase in the last 90 days.",
        "source_of_truth": "billing_db.orders",
    },
    "overdue invoice": {
        "definition": "An invoice where today is later than due date plus grace period.",
        "grace_period_days": 5,
        "source_of_truth": "billing_db.invoices",
    },
    "top client": {
        "definition": "A client in the top 10% by trailing 12-month recognized revenue.",
        "source_of_truth": "finance_dw.revenue_by_client",
    },
}

UNDEFINED_TERMS: list[str] = []  # in production, a ticket queue for the glossary owner


class Run:
    """State of one agent run. Once halted, every tool refuses to work."""

    halted = False
    reason = ""


RUN = Run()  # one run per server for brevity; key this by session in practice


def require_running():
    if RUN.halted:
        raise RuntimeError(f"Run halted ({RUN.reason}). All tool calls are blocked.")


def halt_on_undefined(term: str):
    """The stop is enforced here, in code, not requested in a prompt."""
    UNDEFINED_TERMS.append(term)
    RUN.halted = True
    RUN.reason = f"undefined term '{term}'"
    raise RuntimeError(
        f"'{term}' is not defined. The run is halted and the gap was logged for the glossary owner."
    )


class ClientMeaning(BaseModel):
    meaning: Literal["account", "company"] = Field(
        description="Which meaning of 'client' do you intend?"
    )


@mcp.tool()
async def define_term(term: str, ctx: Context) -> dict:
    """Return the governed business definition of a term."""
    require_running()
    key = term.lower().strip()

    if key in GLOSSARY:
        return GLOSSARY[key]

    if key == "client":
        # Two approved meanings exist. Only the person can say which one this task needs.
        result = await ctx.elicit(
            message="'client' has two approved meanings here. Which one do you intend?",
            schema=ClientMeaning,
        )
        if result.action == "accept":
            return {"term": "client", "meaning": result.data.meaning, "scope": "this task only"}
        raise RuntimeError("'client' is ambiguous and the user did not choose.")

    halt_on_undefined(key)


def run_governed_query(source: str, filters: dict) -> dict:
    ...  # runs the query against the semantic layer


@mcp.tool()
def query_metric(term: str, filters: dict) -> dict:
    """The only data path the agent gets. Approved terms only. There is no raw SQL tool."""
    require_running()
    key = term.lower().strip()
    if key not in GLOSSARY:
        halt_on_undefined(key)  # a synonym or an unknown term halts the run, it is not interpreted
    return run_governed_query(GLOSSARY[key]["source_of_truth"], filters)


class SendApproval(BaseModel):
    approve: bool = Field(description="Send these emails?")


@mcp.tool()
async def confirm_send(recipient_count: int, ctx: Context) -> bool:
    """Ask the person to approve a consequential action before it happens."""
    require_running()
    result = await ctx.elicit(
        message=f"This will email {recipient_count} contacts about overdue invoices. Send?",
        schema=SendApproval,
    )
    return result.action == "accept" and result.data.approve


if __name__ == "__main__":
    mcp.run()
Enter fullscreen mode Exit fullscreen mode

Elicitation and the stop: asking a person only when it matters

Look at what the server does in each case. A defined term returns its governed definition, with no human involved. An undefined term does two things at once: it logs the gap so the glossary owner can fix it, and it halts the run. That log turns "the agent was dumb" into "we found a hole in our definitions."

The halt is enforced in code, not in a prompt. A message that says "do not assume" is only a request, and a model can ignore it by swapping in a synonym that is defined or by writing raw SQL against the tables. So the design removes those options. The agent has no raw SQL tool, query_metric is its only data path and accepts only approved terms, and once the run is halted every tool, including confirm_send, refuses to work. The model can still say whatever it likes, but it has nothing left to call.

Elicitation is used for the two cases that really need a person. The first is ambiguity: "client" has two approved meanings, and only the person knows which one this task needs. The second is a consequential action: before the emails go out, the server pauses and asks for a yes or no. In both cases the server sends the client a small schema describing the answer it wants, and the person can accept, decline, or cancel. The tool handles each outcome.

I deliberately did not use elicitation to let a user type in a missing definition and save it. If it did, whoever happened to be using the agent could redefine "overdue" for the whole company, and nobody would review it. Definitions belong to a governed source. People belong in the loop for ambiguity and approval.

A few details are worth knowing:

Elicitation goes to the user through the client, not to the model. If you want the server to ask the model something, that is a different MCP feature called sampling.
It is an optional client capability, so handle clients that do not support it. For confirm_send, the safe fallback is to refuse to send.
Agents running unattended have nobody to ask. A halted run should hand off to a queue or an approval step for the glossary owner, not wait for an answer that will never come.

Ontology rules as a check before acting

Second, the ontology rules as a check that runs before any email goes out

from datetime import date, timedelta

GRACE_DAYS = 5

def is_overdue(invoice: dict, today: date) -> bool:
    return today > invoice["due_date"] + timedelta(days=GRACE_DAYS)

def can_approve_payment(contact: dict) -> bool:
    return contact["role"] == "Finance"

def validate_followup(invoice: dict, contact: dict, today: date) -> list[str]:
    problems = []
    if not is_overdue(invoice, today):
        problems.append("Invoice is not overdue under the grace-period rule.")
    if contact["customer_id"] != invoice["customer_id"]:
        problems.append("Contact does not belong to the invoiced customer.")
    if not can_approve_payment(contact):
        problems.append("Contact is not in Finance and cannot approve payment.")
    return problems
Enter fullscreen mode Exit fullscreen mode

In an agent workflow this sits as a node between "draft the email" and "send the email." If validate_followup returns anything, the agent either fixes the recipient or stops and asks. In a larger system the rules would live in a graph store or a rules engine instead of hard coded functions, but the idea is the same: the knowledge is written down outside the prompt, and something checks it.

Remove any row and both the person and the agent start to look dumb. Fill them all in and an ordinary worker, human or digital, becomes dependable.

Where to look first when an agent fails

Before you swap the model, check these:

  • Did I define the terms the task depends on?
  • Did I describe how those concepts relate?
  • Are the rules and boundaries written down where something can check them?
  • Is there a feedback loop, such as evals, so the same mistake does not repeat?

Most of the time the answer to one of these is no. That is a design gap on our side, not a flaw in the model.

Takeaway

Calling someone dumb when you never gave them direction says more about the teaching than the learner. Agents are no different. Semantic models and ontologies are how you turn tribal knowledge into shared knowledge, and shared knowledge is what lets any worker do the job the way you meant it.

The next time an agent gets something wrong, ask what it did not know, and then teach it.

Thanks
Sreeni Ramadorai

Top comments (4)

Collapse
 
reidmarlow profile image
Reid Marlow •

The distinction between run-time elicitation and governance updates is the sharp part here. Exposing glossary mutations directly to an agent or end-user invariably turns business metrics into a moving target across departments.

One failure mode I run into with semantic lookup tools is prompt-loop evasion: when a model hits an undefined term error, it often tries to work around the missing definition by substituting synonyms or raw SQL queries rather than stopping. Enforcing the halt at the execution boundary rather than relying on prompt adherence is usually what keeps the guarantee intact.

Collapse
 
sreeni5018 profile image
Seenivasa Ramadurai •

Thanks, this is a good point. The post is a design idea, not a production harness, so the "do not assume" error in my sketch is only a prompt-level instruction, and you're right that a model can work around it with a synonym or raw SQL. The halt has to be enforced at the execution boundary, not left to the prompt. In an actual design that means the agent reaches data only through governed tools, with no raw SQL path, and an undefined-term result puts the run into a stop state it can't leave on its own. The exact mechanism depends on the architecture and the orchestration layer. I'll add a section on this to the post.

Collapse
 
irr123456 profile image
Ivan •

Interesting approach. Have you tried running the same case many times to see how stable the results are?

In my experience non-determinism often stays a bottleneck even with good context.

Collapse
 
sreeni5018 profile image
Seenivasa Ramadurai •

Not yet. This post is a design idea I had recently, and I wanted to share it and hear how others see it, so I haven't run it many times or measured how stable it is. I agree that non-determinism stays a bottleneck even with good context, since context narrows the range of outputs but doesn't remove the variation. The thinking behind the design is to move as much as possible out of the model and into deterministic pieces: governed definitions, rule checks in code, and a halt enforced by the tools. The model's variation then affects the plan, but not the meaning of the terms or the final checks. I'd still want to test that by running the same case many times and seeing how often it stays correct. How do you measure stability in your own runs?