The search over your documents
was built a year ago.
Every page was turned into a vector,
a long list of numbers
that places its meaning
somewhere in space,
and stored in an index.
A question is turned into numbers
the same way,
and the pages nearest to it
come back as the answer.
Then a better embedding model came out.
Cheaper, too.
So new documents
and new questions
started going through the new one.
Nobody re-embedded the old pages.
There were two hundred thousand of them,
and it would have cost a weekend
and a noticeable bill.
Nothing broke.
That is the trouble.
Two models do not share a space.
The numbers one model produces
mean nothing
in the coordinates of the other.
Same length, sometimes.
Same kind of values.
No common map.
So every new question
is measured against the old pages
as if they stood in the same room,
and the distance between them
is noise that looks like a score.
The search still returns five results.
They still carry similarity numbers.
The answers written from them
are still fluent.
They are just built
on whatever happened to land nearby.
And the new documents,
which do share the new space,
win every time,
so the older knowledge
quietly stops being found at all.
Nobody sees an error.
People get vaguer answers
and decide the assistant
has gone off.
Treat the embedding model
as part of the schema.
Store the model name and version
beside every vector.
Refuse to compare vectors
from different versions,
loudly,
in code.
When you change the model,
build a new index from scratch
alongside the old one,
and switch only when
every document is in it.
Keep a small set of questions
with the pages they should find,
fifty is plenty,
and run it against both indexes
before you switch,
so better is a number you measured
and not a line in a release note.
And count the cost of re-embedding
when you choose a model,
not halfway through adopting it.
An index is not a pile of documents.
It is a pile of documents
as one particular model saw them.
Change the eyes,
and you have to look again.
– Serguey Asael Shinder
Top comments (0)