TL;DR — Vector databases let embeddings from different model versions sit in the same index with no complaint, because the vectors are type-valid even when they're geometrically meaningless together. At data-warehouse scale, re-embedding billions of rows after a model upgrade is a massive, slow batch job, which means mixed-version indexes aren't an edge case — they're the default state during every migration. The fix is treating model identity as a schema constraint on the embedding column, not metadata you check if you remember to.
A vector column passes every check your warehouse knows how to run. It's the right dimensionality, the right dtype, non-null, within the expected norm range. Nothing in the schema catches the one thing that actually matters: whether every vector in that column came from the same model. At data-warehouse scale, it usually doesn't, and the failure that follows is one of the quietest in the entire AI stack.
The schema that lies
Embeddings are not data. They're coordinates in a geometry that a specific model defined. Cosine similarity between two vectors is only meaningful if both vectors were placed in the same space by the same function. Swap the model — a new checkpoint, a new tokenizer, even a changed normalization step — and you get a different space with the same shape. A 1536-dimensional float array from one model version looks structurally identical to one from the next version. Your warehouse's type system has no way to tell them apart, because it was never designed to.
This is the gap: relational schemas enforce structural constraints — type, nullability, foreign keys between tables. They have no concept of "produced by model X at checkpoint Y with preprocessing Z." So when a column of embeddings gets partially regenerated, the warehouse sees a clean column of valid vectors. The vector database built on top of it sees a clean index. Nearest-neighbor search runs without a single error. The only symptom is that retrieval quality drops, slowly, in a way that looks like model drift or data quality rot rather than what it actually is — two incompatible coordinate systems being compared as if they were one.
Why this isn't a small-scale problem
At small scale, this issue barely exists. Re-embed ten million rows against a new model, swap the table, done in an afternoon. The transition is close enough to atomic that nobody notices the seam.
At warehouse scale — billions of rows accumulated across years, embedding columns attached to historical fact tables, feature tables, document archives — re-embedding is not an afternoon job. It's a sustained batch workload that can run for days, sometimes longer depending on throughput and cost constraints. During that entire window, the production index contains a blend: rows re-embedded under the new model sitting next to rows still carrying the old vectors, because nobody can afford to take the index offline until the backfill finishes.
This means mixed-version indexes aren't a rare migration accident. They're the default operating condition every time a warehouse-scale system upgrades its embedding model. The bigger the warehouse, the longer that window stays open, and the longer retrieval quality silently degrades without anyone being able to point to a cause, because every individual vector in the index is perfectly valid on its own.
Re-embedding is the hidden line item in your compute budget
There's a cost dimension here that most capacity planning misses. Inference serving cost scales with query volume — you can forecast it, cap it, autoscale around it. Re-embedding cost scales with total historical data volume, which only grows. Every time you adopt a better embedding model, you're not paying a marginal cost on new data; you're paying to reprocess everything that came before it, going back as far as your retention policy allows.
For a warehouse holding years of documents, support tickets, logs, or transcripts, that backfill pass can dwarf the ongoing inference bill. Teams that treat re-embedding as a one-off script run by whoever happens to own the vector database end up re-discovering this cost the hard way, usually mid-migration, usually after the budget was set assuming embedding was a solved, cheap preprocessing step rather than a recurring, data-volume-scaled compute commitment.
The practical consequence is that re-embedding needs to be budgeted and scheduled like any other large backfill: with its own SLA, its own compute allocation, and its own owner — not treated as an incidental side effect of a model upgrade decided in a planning doc somewhere else.
Provenance as a constraint, not a metadata field
The fix isn't exotic. It's the same move relational databases made decades ago when they decided that referential integrity shouldn't be optional: make the dependency explicit and enforce it structurally instead of trusting people to remember it.
Concretely, every embedding column should carry a composite identity alongside the vector — model identifier, checkpoint or version hash, and preprocessing hash — and that identity should be a first-class part of how the vector index is partitioned and queried, not an optional tag in a sidecar table. A similarity search spanning two different identities shouldn't silently succeed; it should be something the system either refuses by default or flags loudly, the same way a join on mismatched types gets rejected rather than silently coerced.
Operationally, this looks like a few concrete habits:
Treat re-embedding as a first-class ETL pipeline with lineage, monitoring, and a defined completion state — not a script someone runs once and forgets.
Build the new index fully, under the new model identity, before cutover — a blue/green swap for vector indexes, the same discipline already standard for schema migrations.
Filter or partition retrieval by model identity during the transition window, so a query never silently blends two incompatible spaces even while the backfill is still running.
Log model identity as part of every embedding's provenance record, queryable the same way you'd query column lineage, so a quality regression can be traced to "which model wrote this vector" in minutes instead of days.
None of this requires new infrastructure. It requires deciding that model identity is a dependency the schema should enforce, not a fact that lives in a changelog nobody reads during an incident.
The real lesson
Vector databases inherited the warehouse's confidence that if the types check out, the data is sound. That confidence does
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.