For years, "AI data" lived somewhere else — a separate vector database bolted onto the side of
your warehouse, kept in sync with brittle pipelines. That's changing. Modern warehouses can
store embeddings as a column and run similarity search in SQL, right next to your structured
data. That unlocks a genuinely new modeling pattern: hybrid queries that filter with
ordinary SQL and rank by semantic similarity in the same statement.
Why keeping vectors in the warehouse matters
One copy of the data, one governance model, one query engine. You can write something like
"find the records in this region, from last quarter most similar to this description" — a
structured filter and a semantic rank together — without shuttling data between systems or
reconciling two sources of truth. No sync pipeline to break at 2am.
The modeling decisions that make or break it
Embeddings are data, so they need the same discipline as any model:
- Pick the embedding grain. What is one vector of? A row? A text field? A record plus its context? This is a grain decision, and it determines whether retrieval returns signal or noise. Too coarse and results are vague; too fine and they lose meaning.
- Version the embeddings. The moment you change embedding models, old and new vectors aren't comparable. Store the model/version alongside each vector so you know what's stale, and plan how you'll re-embed.
- Keep rich metadata. Store source, date, entity, and permissions next to each vector so you can filter retrieval, not rely on similarity alone. Metadata is what makes hybrid search work.
- Mind the cost. Embedding generation and vector indexes aren't free. Model the refresh cadence deliberately — re-embed on change, not on every run.
What it's good for
- Semantic search over your own records — tickets, docs, products — governed like the rest of your data.
- Entity resolution / dedup — near-duplicate records that exact matching misses.
- RAG over the warehouse — serving trusted, filtered context to an LLM straight from governed tables instead of a disconnected store.
A caution
Vector search returns the nearest match, not the correct one — it's confident even when
it's wrong. Treat similarity as a ranking signal, keep a relevance threshold, and always carry
metadata so you can filter and audit what came back. Semantic ≠ accurate.
Takeaways
- Warehouses can now hold embeddings as columns and do similarity search in SQL — one copy, one governance model.
- Hybrid queries (SQL filter + semantic rank) are the new superpower; no separate vector DB to sync.
- Treat embedding grain, versioning, metadata, and refresh cost as first-class modeling choices.
- Nearest isn't correct — threshold and filter with metadata.
Top comments (0)