In the comparison articles (part 1 and part 2) I put my own RAG against five other systems. The models were frozen there on purpose: same embedder,...
For further actions, you may consider blocking this person and/or reporting abuse
Great write-up. The line I'd underline is that the failures no model swap fixed were retrieval-side. Two examples from the live vector · k5 index (single queries, not systematically repeated):
A short verbatim quote isn't answered. Searching a 4-word quote from p. 108 ("The communication shall be") returned an answer saying the question is incomplete, with Evidence: 0 passages and the warning "The model cited no specific passage". The answer still lists topics from the sources it received (confidentiality, entry into force, duration of undertakings, communication channels), so retrieval did return passages, just not necessarily the one with the quote, or the model didn't use it. I can't tell which from the UI. Either way it looks like a limit of pure vector retrieval on fragment queries. A lexical leg would help, but your language article shows the risk: FTS matched 192 of 484 chunks on shared acronyms and the fusion trusted it. Showing retrieved vs cited passages in the UI would tell a retrieval miss from a generator refusal.
Duplicated rule text eats Top-K. "Voting by written procedure" returned identical blocks from sct_inst_rulebook_2025.pdf (p. 108) and sct_rulebook_2025.pdf (p. 106), 2 of 5 slots. I'd collapse duplicates before building the context but keep both sources in the citation, since "this rule applies to both SCT and SCT Inst" is information in itself.
Fusion depends on the query type. In my own ablation (30 code-repository tasks, single run), FTS5 alone scored recall 0.825, the full pipeline without rerank 0.775, and adding the vector tier on top of BM25 lowered recall by 0.098. Recall there is measured as evidence-pattern hits in the returned text. That's an identifier-heavy corpus, so I read it as "the hybrid decision depends on the query type", not as a contradiction of your regulation-text results. Log: exp-5 in the Lab section.
Building on Edward's point: will Phase 2 include dedup and chunk headers, and would Recall@5 by gold_chunk_id show whether the 0-passage case is a ranking problem or a generator problem?
Follow-up to point 1, with a few more manual runs on the same section (single queries, so anecdotes rather than measurements).
"Within how many calendar days can a PSMB member request a telephone meeting?" returned "the sources do not contain information" with 0 cited passages in 2 of 2 runs, and the answer quotes the general meetings rule instead. The 5-day rule sits in the written-procedure paragraph (p. 106/108).
Two rewordings that keep the written-procedure context ("…instead of the written vote?", "…after receiving a written communication with a proposed decision?") both answered 5 calendar days with p. 106 and p. 108 cited. Two other details from the same paragraph (the Secretariat's 2-working-day deadline for communicating results, and the same legal force as a meeting) were also answered correctly.
So the content is retrievable; what changed is whether the query carries the section's topic words. I can't see the retrieved chunks, but the failing query looks like a retrieval miss that the answer prompt turned into a false refusal. It might be worth adding "answer exists in the corpus but the question lacks section context" cases next to your 20 absent-answer traps.
I reproduced your runs against the stored production vectors. Your reading was right on all three points.
The telephone-meeting question: the passage that holds the 5-day rule was at rank 7 for the plain vector leg, so the refusal was a retrieval miss, not a generator decision. The top 5 were the general PSMB meeting pages, and 4 of those 5 slots were two duplicate pairs, so the generator saw three distinct passages. Your two rewordings put the same passage at rank 1.
Duplicates: 138 of the 484 chunks (28.5%) are byte-identical across the two rulebooks, because the 36-page EPC Payment Scheme Management Rules are bound into both. Collapsing exact duplicates before building the context moves the 5-day passage from rank 7 to rank 4 on your original question. Adding the full-text leg moves it to rank 3 (full-text alone had it at rank 1). Either of your fixes alone would have answered it.
The 4-word quote: vector had the passage at rank 25, full-text at rank 1. Fragment queries are a lexical job, agreed, and the fusion problem from the language article stays; the compromise I am testing is fusing only when the query language matches the index.
Retrieved vs cited in the UI: yes, and it would have told you the answer to point 1 in one look. Going on the list. Publishing soon the full numbers
Thanks for actually re-running all three and giving the ranks back. That's more than most people do.
Your two fixes do different jobs, though. Dedup frees the slots (7→4), the full-text leg fixes the quote (25 → 1). Not either/or — they stack.
One thing to watch: you said you're testing fusion only when the query language matches the index. Acronyms don't change with language, so that rule may not catch them. In the language article FTS matched 192 of 484 chunks on shared acronyms. Does your test cover that case?
And I'd still push one more trap. The answer is in the corpus, the question just doesn't have the section's topic words — my telephone-meeting one refuses 2 out of 2, while two rewordings answer it. That's retrieval losing to phrasing, not a generation problem. Would you add it next to your 20 absent-answer traps?
Agreed, they stack. On your telephone question either fix alone reached the top 5: dedup rank 4, full-text rank 3. On the 4-word quote dedup left it at rank 25 and only full-text moved it. Dedup frees the slots, full-text matches the words in the question. The follow-up goes out this week with both numbers.
On the language rule I overstated it. It is argued in the language article and not built. As written it uses the language of the whole query. "SCT Inst" inside a Russian question still counts as Russian. It fails where you point: acronym-only queries, and English questions, where full-text still cost 2 to 3 questions of 20. The same article names the fix for that: a chunk counts only when more than one term matches.
Yes to the extra trap. My 20 all test one thing: the answer is not in the corpus and the model should say so. Yours tests the opposite: the answer is there, but the question does not have the section's topic words. The judge score cannot see that. The rank of the right passage can. Your telephone question is the first one I add.
We saw the same embedder gap. We started on a mixed Korean-English financial corpus with a general multilingual model, and it honestly just couldn't produce useful rankings on Korean queries. Generator swaps moved things a couple points; swapping the embedder was closer to 8 points on Korean specifically, with English basically flat. The refusal thing we'd already discovered the hard way: upgraded a generator expecting fewer hallucinations on out-of-scope questions, rates didn't move at all because the boundary was in the prompt the whole time.
Your Korean-English gap is the reason I went and measured it. Same shape on my side: Spanish and Russian questions against the English SEPA corpus cost text-embedding-3-small 2 to 3 questions in 20, and the Gemini and Voyage embedders lose about the same. The surprise was elsewhere: hybrid search lost another 2 to 3 in every language, English included, because the full-text leg matched 192 of 484 chunks on shared acronyms and the fusion trusted it. Writing it up now, will link it here.
Edit: written here Asking a RAG in the wrong language
we got bitten by the exact failure your matrix describes, except ours was
quieter and took a week to catch. last february we swapped from
text-embedding-3-small to bge-large thinking "better embedder, better RAG."
golden set stayed green, two negative tests that had been failing for months
suddenly passed. looked like a win until a real user question hit the same
trap from a different angle — the new embedder just stopped retrieving the
trap chunks at all, so the model "passed" by luck, never seeing the bait.
your 3.8-point embedder delta matches what we saw: the same question, same
generator, same corpus, different embedder = completely different retrieval
behavior. our fix was tagging every DECISIONS.md line that makes a retrieval
claim with the embedder it was earned under. after any swap, lines with the
old tag are stale by default — they stay in the file, but CI flags them as
"unverified under current embedder" until someone re-runs and re-dates them.
cheap bookkeeping, stopped us from trusting verdicts that describe a
retrieval that no longer exists.
the trap_chunk_id thing from our chat with edward izgorodin fits here too:
every negative test carries the id of the chunk containing the trap, and CI
fails if that chunk isn't in the retrieved set, regardless of what the answer
says. "test did not run" becomes a third verdict next to pass and fail.
maintenance cost is real though — trap ids move whenever we re-chunk, so the
stamping step has to re-run after any pipeline change, same as the embedder
tag. two stamps per line now.
your finding that refusal discipline lives in the prompt, not the model, is
the sharpest takeaway. we learned that the hard way too: spent two days
debugging why one generator hallucinated on traps while another refused,
turned out the prompt was different between the two test harnesses. same
prompt, same refusal rate across all six generators you tested. boring
conclusion, but it's the one that actually stuck.
The 'negative tests suddenly passed' story is the one I would put next to 'no model purchase fixes a corpus property'. Your tagging plus trap-chunk setup is more rigorous than my refusal traps; if you have written it up anywhere I would link it.
not written up anywhere yet, which is honestly a gap — the whole setup lives
in our DECISIONS.md and CI config, not in public. your comment just moved it
to the top of my writing list: drafting a post this week with the exact
mechanics (trap_chunk_id stamping at authoring time, embedder tags per
retrieval claim, the "test did not run" third verdict, and the two-stamps-
per-line maintenance cost nobody warns you about). will drop the link here
the moment it's up, and then it's your call whether it deserves the link :)
writeup is live — the comment spam filter eats raw links from my account, so:
first post on my profile, title "The Negative Test That Passed for the Wrong
Reason."
went with the full mechanics: trap_chunk_id stamping at authoring time,
embedder tags per retrieval claim, the "test did not run" third verdict, and
a section on the re-stamping cost after every re-chunk. your "corpus decides
more than the model" line is in there as the framing for why the embedder
slot deserves its own regression test. thanks again for pushing on it — the
post exists because you asked for it.
This is the experiment most RAG writeups skip: holding the pipeline fixed and moving only the models. Two results stand out to me. The embedder moving scores more than the generator matches what I have seen, and it is the slot nobody benchmarks. And "no model purchase fixes a corpus property" is the sentence I would put at the top of every retrieval postmortem.
The refusal result is the one I want to push on. All six generators refusing all twenty traps, including the cheapest, says the prompt boundary is doing the work, not the model. So my question is what happens when the trap is adjacent rather than absent: a question whose answer is in the corpus but under a conflicting passage. That is where I have seen cheap generators start blending sources, and it would be interesting whether your matrix separates pipelines there too.
The adjacent trap is a good one and I have a first data point, from the wrong direction. The refusal boundary held on all 20 traps in English, and on 14 of 15 cross-lingual runs of the trap questions; the 15th, in Russian, answered a direct-debit question with the credit-transfer recall period from an adjacent passage. So a conflicting-but-present passage does get blended, at least by the cheapest generator, when the question is far enough from the corpus language. A proper adjacent-trap set is on the list.
The line that the corpus decides more than the model does is the result I would build on, and the 3 questions that fail on all six pipelines are probably the most informative rows in the matrix. If the golden set records which passage holds each reference answer, those three split into two different problems: the passage never reaches the top five, which is a chunking or embedding question no generator can fix, or it reaches the top five and every generator still misreads it, which points at the passage itself. The same split would sharpen the 11 questions that depend on model choice, since an embedder swap that changes the answer should also change whether the passage was retrieved, and a generator swap should not.
It also gives the embedder comparison a number that needs no judge at all. Whether the passage was in the top five is a lookup, not a verdict, so it cannot inherit a judge's preference for its own family.
I ran the split you proposed. Retrieval in the matrix is decided by the embedder alone (four of the six pipelines share text-embedding-3-small and return the same five passages), so the 14 non-unanimous questions give 14 x 3 retrieval cells. "Retrieved" means a phrase from the reference answer occurs in one of the five passages, a regex, no model.
Of the 3 questions that fail on all six pipelines, 2 were never retrieved by any embedder (the BSD-3 clause, whose license name exists only in the filename, and the SCT Return deadline). The third (MPL Larger Work) had the definition in the top 5 on two embedders but not the permission clause, and both generators refused.
Of the 11 model-dependent questions, 4 follow retrieval exactly: the passage is absent on the embedders that scored 0 and present on the ones that scored 100 (LangGraph Send, Apache patent retaliation, the saturated fat limit, the SCT execution time). 6 had the passage at rank 1 or 2 on every embedder and still split, so those are the generator or the judge threshold. And 1 is the case you warned about: the Apache patent grant was never retrieved on any embedder, GPT answered "Yes, though the provided excerpt does not contain its specific terms", and the judge gave it 100 with grounded=false. The lookup catches it; the verdict did not.
Written up with the tables, publishing soon.
The Apache row is the one I would build on next, because on it two instruments that share nothing agreed. The lookup found no passage, and the same judge that scored the answer 100 also returned grounded=false. A string search and a model reading a model can each be wrong, but not in the same way.
That is why I would run the anchors over the 66 questions that passed on all six pipelines, which the split has not touched yet. A unanimous pass says nothing about where the answer came from, and the matrix already had not-grounded verdicts scoring 70 or above. A pass with the anchor absent is the Apache case again, in a place nobody has looked. A disagreement between the anchor and the grounded flag is a bad anchor, a judge misread or an answer that added something the passages do not hold, and the first two have both happened here already: two anchors corrected by hand, and the groundedness judge flagging Sources lines before its input was fixed.