The model calls were expensive,
so somebody added a cache.
Not an exact match cache.
People never ask
the same question the same way,
so the cache turned each question
into an embedding,
and if a new question was close enough
to an old one,
it served the old answer.
Close enough
meant a similarity above 0.92.
The hit rate was wonderful.
The bill dropped by half.
Then a customer asked
how to stop her order
being cancelled.
She was told
how to cancel her order.
Instantly.
From the cache.
The two questions
used almost the same words,
about the same thing,
in the same tone.
To an embedding
that is what similar means.
It measures what a sentence is about.
It is very bad at what a sentence says.
Not,
before and after,
refund and charge,
this account and that one,
are small words
inside a large shared topic,
and they barely move the vector.
They are also the whole meaning.
A cache built this way
does not save you model calls.
It replaces some of them
with a confident guess
about which earlier person
asked your question.
So decide what you are allowed to reuse.
Static explanations,
where the right answer
is the same for everyone,
can be shared.
Anything that leads to an action,
or depends on whose account it is,
or touches money,
goes to the model every time,
or to a cache keyed exactly
on the normalised question,
the user,
and the version of the documents
the answer came from.
When the help page changes,
everything built on it expires.
Then test the threshold
the way an attacker would.
Write pairs that share words
and mean opposite things.
Cancel and keep.
Include and exclude.
Before the deadline and after it.
Run them against the cache
and count how many come back
as the same question.
That number should be zero.
And log every cache hit
with both questions side by side,
so somebody can read
what the cache decided
was close enough.
Similar is a statement about topics.
Answers are about what was asked.
– Serguey Asael Shinder
Top comments (0)