Three embedding launches landed in one week.
Every one came with a chart.
None of the charts were drawn on your documents.
- On September 24, 2026, TopK introduced topk-embed-v1, a family of multi-vector embedding models for text and visual documents. Two smaller sizes are open source.
- On September 30, Cohere released Embed 5 in Pro and Fast tiers. "Both tiers share a single embedding space, so teams can index with Pro and query with either model without rebuilding the index." Across 40 development datasets, Cohere says the mixed pairs averaged 1.6% and 2.7% losses against same-model baselines.
- The same day, Perplexity introduced pplx-embed-v2-context-9b-preview, trained "to retrieve both answer chunks and supporting context rather than a single gold chunk." It was scored on context-bench, which turbopuffer keeps private "to reduce the risk of training contamination."
Different products. Same direction.
Retrieval is no longer one number.
Did you find the answer? Did you find what proves it? Can you swap the query model without rebuilding the index?
The Perplexity model card has my favorite warning of the week. Embeddings from the preview "should not be mixed with embeddings from a future release."
That's the part I keep thinking about.
Shared space is a claim. Mixing is a migration.
The only benchmark that knows your documents is the one you build.
So let's build a tiny one.
By the end, you'll run one command:
npx tsx eval.ts
And watch an eval compare two embedding models, decide whether to ship, and catch an index you should not mix.
No API key.
No model.
Just TypeScript.
One honesty note: the "embedding models" here are hashed bag-of-words vectors, not neural models. They are stand-ins so the eval has something to grade. Swap in your provider's embed call and keep the rest.
Table of Contents
- What We Are Building
- Project Setup
- Step 1: Label a Tiny Corpus
- Step 2: Mock Embedding Models
- Step 3: Index and Rank
- Step 4: The Metrics
- Step 5: Compare, Gate, and Mix
- Where It Breaks Down
- The Bigger Idea
Code: github.com/bobbyhalljr/tiny-retrieval-eval
What We Are Building
Twelve chunks from two near-identical leases and a deploy runbook. Four labeled questions.
Two models. chunk-only embeds each chunk by itself. That's "prod". contextual adds the document title. That's the candidate.
Then a ship gate, and a mixed-index test for when the query model and the index model disagree.
It's also a small version of the question behind Helix: when the system changes, can you show what got better, what got worse, and why?
Project Setup
You will need Node.js 18 or newer.
mkdir tiny-retrieval-eval
cd tiny-retrieval-eval
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as eval.ts.
Step 1: Label a Tiny Corpus
// eval.ts: a tiny retrieval eval to run before you switch embedding models.
// Everything is mocked: the "embedding models" are hashed bag-of-words vectors, not real models.
// The corpus, queries, labels and the ship tolerance are example inputs. No API key, no network.
// Step 1: a tiny labeled corpus
type Chunk = { id: string; title: string; text: string };
type Query = { id: string; text: string; answer: string; evidence: string[] };
const DOCS: Record<string, [string, string[]]> = {
oak: ["Lease for 12 Oak Street", [
"This lease covers 12 Oak Street, unit 3.",
"Monthly rent is 2,100 dollars.",
"Late fee: 50 dollars after five days.",
"Dogs and cats are allowed with a pet deposit.",
]],
elm: ["Lease for 48 Elm Avenue", [
"This lease covers 48 Elm Avenue, unit 9.",
"Monthly rent is 2,650 dollars.",
"Late fee: 75 dollars after three days.",
"No pets are allowed.",
]],
runbook: ["Payments deploy runbook", [
"Every deploy ships a tagged release.",
"If errors spike after a deploy, page on-call.",
"To roll back, redeploy the previous release tag.",
"Watch errors for ten minutes to confirm the rollback.",
]],
};
const CHUNKS: Chunk[] = Object.entries(DOCS).flatMap(([doc, [title, texts]]) =>
texts.map((text, i) => ({ id: `${doc}#${i}`, title, text })),
);
// answer = the chunk that holds the answer. evidence = every chunk you need to trust it.
const QUERIES: Query[] = [
{ id: "rent at 48 Elm", text: "Monthly rent at 48 Elm Avenue?", answer: "elm#1", evidence: ["elm#1", "elm#0"] },
{ id: "rent at 12 Oak", text: "Monthly rent at 12 Oak Street?", answer: "oak#1", evidence: ["oak#1", "oak#0"] },
{ id: "late fee at Elm", text: "Late fee on the Elm Avenue lease?", answer: "elm#2", evidence: ["elm#2", "elm#0"] },
{ id: "payments rollback", text: "How do I roll back a payments deploy?", answer: "runbook#2", evidence: ["runbook#2", "runbook#3"] },
];
The leases are near duplicates on purpose.
"Monthly rent is 2,100 dollars" and "Monthly rent is 2,650 dollars" look almost the same to a retriever. Only the rest of the document tells them apart.
answer is the chunk that holds the answer. evidence is everything a reader needs to trust it, like the chunk that says which apartment this is.
A gold chunk tells you what to find. Evidence tells you what to show.
Step 2: Mock Embedding Models
// Step 2: mock embedding models (signed feature hashing: same seed, same space)
type Model = { name: string; seed: number; dims: number; useTitle: boolean; dropStopwords: boolean };
const STOP = new Set("a an the is are of on at to for and do i how".split(" "));
function tokens(text: string, dropStopwords: boolean): string[] {
return text.toLowerCase().split(/[^a-z0-9]+/)
.filter((t) => t && !(dropStopwords && STOP.has(t)))
.map((t) => (t.length > 3 && t.endsWith("s") ? t.slice(0, -1) : t));
}
function embed(m: Model, text: string): number[] {
const v = new Array<number>(m.dims).fill(0);
for (const t of tokens(text, m.dropStopwords)) {
let h = 2166136261 ^ m.seed; // FNV-1a
for (const ch of t) h = Math.imul(h ^ ch.charCodeAt(0), 16777619);
h >>>= 0;
v[h % m.dims] += h & 1 ? 1 : -1;
}
const norm = Math.hypot(...v) || 1;
return v.map((x) => x / norm);
}
const chunkOnly: Model = { name: "chunk-only", seed: 1, dims: 256, useTitle: false, dropStopwords: true };
const contextual: Model = { name: "contextual", seed: 7, dims: 256, useTitle: true, dropStopwords: true };
const ctxFast: Model = { ...contextual, name: "ctx-fast", dropStopwords: false }; // same seed, cheaper
const ctxV2: Model = { ...contextual, name: "ctx-v2", seed: 99 }; // next release, new seed
embed hashes each token into one of 256 slots with a plus or minus sign. Same seed, same space. New seed, new space.
contextual sees the document title too. Perplexity's model encodes a document's chunks together, so "each chunk's embedding reflects its surrounding context." Prepending the title is the cheapest version of that idea.
ctx-fast is the cheaper sibling in the same space. ctx-v2 is next year's release.
Step 3: Index and Rank
// Step 3: index and rank
type Index = { model: Model; vectors: Map<string, number[]> };
function buildIndex(m: Model): Index {
const vectors = new Map(CHUNKS.map((c) => [c.id, embed(m, m.useTitle ? `${c.title}. ${c.text}` : c.text)]));
return { model: m, vectors };
}
function rank(index: Index, queryModel: Model, text: string): string[] {
if (queryModel.dims !== index.model.dims) throw new Error("dims mismatch: re-embed the index");
const q = embed(queryModel, text);
return [...index.vectors.entries()]
.map(([id, v]) => ({ id, score: v.reduce((s, x, i) => s + x * q[i], 0) }))
.sort((a, b) => b.score - a.score || a.id.localeCompare(b.id))
.map((r) => r.id);
}
The index remembers which model built it.
rank refuses a query model with different dimensions. Cohere's footnote says it plainly: "Both sides must use the same output dimension."
What it can't check is whether two models with the same dimensions share a space. Same shape is not the same meaning.
Step 4: The Metrics
// Step 4: the metrics
const K = 2;
type Row = { id: string; answerRank: number; hits: number; total: number };
type Score = { answer: number; evidence: number; mrr: number };
function runEval(index: Index, queryModel: Model): { rows: Row[]; score: Score } {
const rows = QUERIES.map((q) => {
const ranked = rank(index, queryModel, q.text);
const hits = q.evidence.filter((e) => ranked.slice(0, K).includes(e)).length;
return { id: q.id, answerRank: ranked.indexOf(q.answer) + 1, hits, total: q.evidence.length };
});
const mean = (f: (r: Row) => number) => rows.reduce((s, r) => s + f(r), 0) / rows.length;
return {
rows,
score: {
answer: mean((r) => (r.answerRank <= K ? 1 : 0)), // answer chunk in top k
evidence: mean((r) => (r.hits === r.total ? 1 : 0)), // every evidence chunk in top k
mrr: mean((r) => 1 / r.answerRank),
},
};
}
Three numbers per model:
-
answer@2: the answer chunk is in the top 2. -
evidence@2: every evidence chunk is in the top 2. -
MRR: one over the answer's rank, averaged.
k is 2 because small context windows are where retrieval choices hurt most. On twelve chunks, k=10 is not an eval. It's a participation trophy.
Step 5: Compare, Gate, and Mix
// Step 5: compare, gate, and test a mixed index
const prod = runEval(buildIndex(chunkOnly), chunkOnly);
const cand = runEval(buildIndex(contextual), contextual);
const keys: (keyof Score)[] = ["answer", "evidence", "mrr"];
const fmt = (s: Score) => keys.map((k) => `${k === "mrr" ? "MRR" : `${k}@${K}`} ${s[k].toFixed(2)}`).join(" ");
console.log(`${CHUNKS.length} chunks, ${QUERIES.length} labeled queries (example data), k=${K}\n`);
console.log(`== switch test ==\nchunk-only ${fmt(prod.score)}\ncontextual ${fmt(cand.score)}`);
const worse = prod.rows.filter((p, i) => cand.rows[i].answerRank > p.answerRank).map((p) => p.id);
prod.rows.forEach((p, i) => console.log(` ${p.id.padEnd(18)} answer rank ${p.answerRank} -> ${cand.rows[i].answerRank}`));
const TOLERANCE = 0.05; // example threshold
const drops = keys.filter((k) => cand.score[k] < prod.score[k] - TOLERANCE);
console.log(`gate: ${drops.length || worse.length ? "HOLD" : "SHIP"} (drops: ${drops.join(", ") || "none"}; worse queries: ${worse.join(", ") || "none"})\n`);
console.log("== mixed-index test (index built with contextual) ==");
const ctxIndex = buildIndex(contextual);
for (const m of [contextual, ctxFast, ctxV2]) console.log(`query with ${m.name.padEnd(11)} ${fmt(runEval(ctxIndex, m).score)}`);
The gate has two rules. No metric drops by more than the tolerance. No single query gets a worse answer rank.
The second rule matters more than it looks. An average can go up while the one query your support team cares about goes down.
Run it:
npx tsx eval.ts
You should see:
12 chunks, 4 labeled queries (example data), k=2
== switch test ==
chunk-only answer@2 0.75 evidence@2 0.50 MRR 0.58
contextual answer@2 1.00 evidence@2 0.75 MRR 1.00
rent at 48 Elm answer rank 2 -> 1
rent at 12 Oak answer rank 3 -> 1
late fee at Elm answer rank 2 -> 1
payments rollback answer rank 1 -> 1
gate: SHIP (drops: none; worse queries: none)
== mixed-index test (index built with contextual) ==
query with contextual answer@2 1.00 evidence@2 0.75 MRR 1.00
query with ctx-fast answer@2 1.00 evidence@2 0.75 MRR 1.00
query with ctx-v2 answer@2 0.00 evidence@2 0.00 MRR 0.21
For "rent at 12 Oak", chunk-only ranked the Elm Avenue rent above the Oak Street rent. A chunk that only says "Monthly rent is 2,100 dollars" doesn't know which lease it lives in.
The contextual model put the answer first on every question.
Answer recall went from 0.75 to 1.00. Evidence recall went from 0.50 to 0.75. No query got worse. The gate says SHIP.
In the mixed-index test, ctx-fast queried the contextual index and scored the same as the original. The shared space held.
ctx-v2 scored zero on both recall metrics. Its MRR of 0.21 is close to what random order gets you on twelve chunks.
Same recipe. Same dimensions. New seed.
Every vector still had 256 numbers. None of them meant the same thing.
That's the failure a dimension check cannot catch. Only an eval can.
Re-embed the index when the space changes. Prove it when someone says it didn't.
Where It Breaks Down
Hashed Words Are Not Embeddings
Look at "payments rollback". Neither model found the confirm chunk, because it says "rollback" and the query says "roll back." A real embedding model handles that. Mine is here to give the eval something to grade, not to win.
Labels Are the Expensive Part
Four queries is a smoke test. One query flipping moves answer@2 by 0.25. Start with real questions from your logs, and add one every time retrieval surprises you. Perplexity trained on chunk relevance scores from a teacher model. You start with hand labels.
Your Eval Set Can Leak Too
turbopuffer keeps context-bench private so models can't train on it. Your version is sneakier. Tune chunk sizes and prompts against the queries you grade with, and you're overfitting to your own test. Keep a held-out slice.
Rerankers and Re-chunking Move the Target
Cohere's footnote says its RCP-nDCG@10 scores "reflect reranking quality rather than first-stage retrieval performance." If you run a reranker, eval both stages. And labels point at chunk ids. Change your chunker and elm#1 might not exist anymore. Past a toy, label spans of text.
The Bigger Idea
My contract tests post said a model upgrade is a breaking change.
An embedding upgrade is the same thing, with one extra trap. The old vectors are still sitting in your index.
New embedding model
↓
Labeled queries ──→ answer@k, evidence@k, MRR
↓
Gate ──→ no metric drops, no query gets worse
↓
Mixed-index test ──→ can the old vectors stay?
↓
Ship, or re-embed first
The vendor chart provides the shortlist.
The labels provide the definition of relevant.
The metrics provide the comparison.
The gate provides the decision.
The mixed-index test provides the migration plan.
Don't switch embedding models on vibes. Switch on your own queries.
Software should explain itself.
I'm building Helix around this idea: every change should come with the why. What changed, what got better, what got worse, and the evidence behind it.
What critical engineering knowledge is your team losing right now?


Top comments (1)
The ctx-v2 result is the one I'd turn into a guard, not only an eval finding. Same dimensions, new seed, zero recall, and nothing in the vectors themselves says they live in a different space.
So I'd stamp every index (ideally every vector batch) with the embedding model id and version, and have the query path refuse to search when the query model isn't on that index's allowed list. Fail closed, loudly, before ranking.
Your mixed-index test then becomes how a model pair earns its place on that list. "Shared space is a claim" gets proven once per release, and the stamp stops anyone from skipping the proof on a Friday deploy.