DEV Community

Cover image for Modeling Your Warehouse So AI Can Actually Use It
Vaishnav Prabhu
Vaishnav Prabhu

Posted on

Modeling Your Warehouse So AI Can Actually Use It

Everyone wants to point an LLM at their data — "ask questions in plain English," "let
the agent pull the numbers." Then it confidently returns the wrong revenue figure, and
trust evaporates. The problem usually isn't the model. It's the data layer underneath
it. An LLM is only as good as the structure, definitions, and correctness of what it
reads. Here's how to model for that.

Give the model a contract, not a guess

If you let a text-to-SQL tool or an agent freelance joins against raw tables, it will
invent a definition of "active customer" that doesn't match anyone else's. A
semantic layer (or a set of certified metric models) fixes this: define each
metric once — its grain, its filters, its joins — so the model composes from trusted
building blocks instead of reconstructing business logic from column names. The
semantic layer becomes the contract between your warehouse and anything, human or AI,
that queries it.

Make tables legible

Models reason over your schema the way a new analyst would — from names and docs. You
reduce hallucinated joins dramatically just by making the warehouse legible:

  • Clear names. gross_sales beats amt2. The model can't infer what it can't read.
  • Documented grain and keys. State the primary key and the grain of every model; it's what keeps an agent from double-counting.
  • Expose curated views, not raw tables. Point AI consumers at a clean, certified layer so they can't wander into staging and half-built tables.

Treat embeddings as a modeling decision

Retrieval-augmented generation (RAG) lives or dies on what you chose as a chunk.
That's a grain question, same as any fact table. Before you store vector columns next
to your rows, decide: what is one "document"? A row? A paragraph? A record plus its
context? Chunk too coarse and retrieval returns noise; too fine and it loses meaning.
Keep rich metadata alongside each embedding (source, date, entity, permissions) so
you can filter retrieval instead of hoping similarity alone is enough.

Correctness matters more with AI, not less

A dashboard with a bad number invites a double-take. An agent with a bad number states
it as fact and moves on. So the fundamentals you already know carry extra weight:

  • Uniqueness and grain — a fan-out that doubles a metric will be surfaced by an agent with total confidence.
  • Freshness — an agent won't notice the feed is a day stale; your contract has to.
  • Definitions — if "revenue" is ambiguous in the warehouse, it'll be ambiguous (and wrong) in the answer.

In other words: the data-quality gates and reconciliation checks you build for humans
become the guardrails that keep AI honest.

Takeaways

  • Put a semantic/metrics layer between your warehouse and any LLM so definitions are a contract, not a guess.
  • Make the schema legible — clear names, documented grain and keys, curated views over raw.
  • Treat RAG chunking and embeddings as a grain-and-metadata modeling problem.
  • Correctness (uniqueness, freshness, clear definitions) matters more with AI, because an agent repeats your mistakes with confidence.

Related in In the Pipeline*: running those models safely once the data is ready.*

Top comments (0)