<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Serhiy Kucherenko</title>
    <description>The latest articles on DEV Community by Serhiy Kucherenko (@skucherenko).</description>
    <link>https://dev.to/skucherenko</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4048232%2F93791092-fa66-4f4d-be33-57839c5696df.jpg</url>
      <title>DEV Community: Serhiy Kucherenko</title>
      <link>https://dev.to/skucherenko</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/skucherenko"/>
    <language>en</language>
    <item>
      <title>The same question, asked three ways</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Fri, 02 Oct 2026 17:10:02 +0000</pubDate>
      <link>https://dev.to/skucherenko/the-same-question-asked-three-ways-1c8c</link>
      <guid>https://dev.to/skucherenko/the-same-question-asked-three-ways-1c8c</guid>
      <description>&lt;p&gt;A RAG system answers a question in two steps. A retriever turns the question into a vector, compares it with the vectors of every passage in the documents, and hands the five closest passages to a language model. The model writes the answer from those five. If the passage that holds the answer is not among the five, the model cannot answer, however good it is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/skucherenko/the-answer-was-at-rank-7-2o81"&gt;In the previous piece&lt;/a&gt; I measured that second condition directly and found a question whose answer sat at rank 7. The model only ever saw ranks 1 to 5, so it refused. A reader, Mikhail, had found that question by hand and then found that two rewordings of it, both containing the words of the section that holds the rule, put the same passage at rank 1. He asked for a new kind of test question: the answer is in the documents, but the question does not use the section's own words. This piece builds that test and runs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The earlier experiment, in one paragraph
&lt;/h2&gt;

&lt;p&gt;The documents are the two SEPA credit transfer rulebooks from the European Payments Council, 484 passages of about 300 words each, stored as vectors from OpenAI's &lt;code&gt;text-embedding-3-small&lt;/code&gt;. The test set is 20 questions whose answer page I labelled by reading the PDFs, plus Mikhail's. The live system retrieves the five closest passages by cosine similarity, nothing else. Two variants exist in the code but are not switched on: collapsing byte-identical passages before taking the top five (138 of the 484 passages are exact duplicates, because a 36-page annex is bound into both rulebooks), and a hybrid path that also runs a &lt;code&gt;Postgres&lt;/code&gt; full-text search and merges the two rankings.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["one question&amp;lt;br/&amp;gt;3 phrasings"] --&amp;gt; E["embedder&amp;lt;br/&amp;gt;text-embedding-3-small"]
    E --&amp;gt; R["ranking of all 484 passages&amp;lt;br/&amp;gt;by cosine"]
    R --&amp;gt; M["where is the labelled page?&amp;lt;br/&amp;gt;rank 1 .. 484"]
    Q -. "words only" .-&amp;gt; F["full-text search&amp;lt;br/&amp;gt;Postgres, english"]
    F --&amp;gt; R2["ranking by ts_rank"]
    R2 --&amp;gt; M
    R --&amp;gt; D["drop exact duplicates"] --&amp;gt; M&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: the measurement is the rank of the labelled page in the full ranking, not a pass or fail from a judge. Rank 5 or better means the model would have seen the passage. Rank 6 or worse means it would not, whatever it then answered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;p&gt;Three phrasings per question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;original&lt;/code&gt;: the question as it stands in the golden set, written by someone who knows the rulebook, so it shares the section's vocabulary.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;plain&lt;/code&gt;: the same need as a customer or a new joiner would type it. "Bank" instead of PSP, "instant transfer" instead of SCT Inst, and none of the terms the section is built on. Each question has a list of forbidden terms and the runner refuses to start if the plain phrasing contains one.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;terse&lt;/code&gt;: three to five words, what a person types into a search box. It may reuse section terms; here the variable is length.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Original&lt;/th&gt;
&lt;th&gt;Plain&lt;/th&gt;
&lt;th&gt;Terse&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;How long does the Beneficiary PSP have to send a Return in a standard SCT?&lt;/td&gt;
&lt;td&gt;The account I paid to is closed. How many days does the receiving bank have to send the money back to my bank?&lt;/td&gt;
&lt;td&gt;closed account money sent back when&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does the SCT scheme itself set a maximum transaction amount?&lt;/td&gt;
&lt;td&gt;Is there a cap on how much I can send in one SEPA transfer?&lt;/td&gt;
&lt;td&gt;SEPA transfer cap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Within how many calendar days can a PSMB member request a telephone meeting instead of the written vote? (Mikhail's rewording)&lt;/td&gt;
&lt;td&gt;Within how many calendar days can a PSMB member request a telephone meeting? (his first question)&lt;/td&gt;
&lt;td&gt;PSMB telephone meeting days&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What I expected:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;H1. The plain wording pushes several answers out of the top five, and the terse wording more. Four questions already sat at rank 4 or 5 with their original wording, so at least those are exposed.&lt;/li&gt;
&lt;li&gt;H2. Dropping duplicates helps the plain and terse phrasings more than the originals, because it frees slots that duplicate passages otherwise take.&lt;/li&gt;
&lt;li&gt;H3. In the language piece the hybrid path cost 3 questions in 20 on English questions, and I blamed passages matching on a single shared acronym. A rule that only counts a passage when it matches more than one query word should recover most of that.&lt;/li&gt;
&lt;li&gt;H4. A rule that switches the lexical leg off when the query language differs from the corpus changes nothing here by construction; every phrasing is English.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first held. The second and third did not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;

&lt;p&gt;Everything runs against the stored production vectors over a read-only connection. Each phrasing is embedded once, compared with all 484 passage vectors, and the rank of the first labelled page is recorded. The full-text leg is the production query: &lt;code&gt;plainto_tsquery&lt;/code&gt; with its terms joined by OR, ranked by &lt;code&gt;ts_rank&lt;/code&gt;. The hybrid leg merges the top 20 of each list with reciprocal rank fusion, as the code does. Two extra legs implement the H3 rule with a threshold of 2 and of 3 distinct matched query words.&lt;/p&gt;

&lt;p&gt;Mikhail's question is the twenty-first row. Its &lt;code&gt;original&lt;/code&gt; is his rewording that answered on the live demo, its &lt;code&gt;plain&lt;/code&gt; is the question he actually asked first and was refused.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: I wrote the plain and terse phrasings before the run and read them back against the labelled pages. Golden sets are the ground truth here, so changing one after seeing the ranks is not allowed; these are the phrasings the run used.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Phrasing alone, vector retrieval
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phrasing&lt;/th&gt;
&lt;th&gt;Answer in top 5&lt;/th&gt;
&lt;th&gt;of 21&lt;/th&gt;
&lt;th&gt;Mean reciprocal rank&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;original&lt;/td&gt;
&lt;td&gt;0.81&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;0.61&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plain&lt;/td&gt;
&lt;td&gt;0.67&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;0.38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;terse&lt;/td&gt;
&lt;td&gt;0.43&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;0.27&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The plain wording loses 3 of the 17 hits and gains none: the weekend-availability question falls from rank 2 to 46, the value-limits question from 4 to 8, and Mikhail's from 1 to 7. The terse wording loses 9 and gains 1. Eight of the 21 questions stay in the top five under all three phrasings, three are outside it under all three, and the remaining ten depend on the wording.&lt;/p&gt;

&lt;p&gt;The plain wording is not uniformly worse, it is different. It moves the standard execution-time question from rank 5 to 2, the recall-deadlines question from 4 to 1, and the charging-principle question, a miss at 38 with the original wording, to 15. The embedder does not prefer the rulebook's vocabulary; it prefers whatever the passage happens to say, and the passage on Recall deadlines happens to say "Duplicate", "Fraudulent" and "sent".&lt;/p&gt;

&lt;p&gt;The weekend question is the clearest loss. "Is the SCT Inst scheme available on weekends and public holidays?" finds the page that says "available 24 hours a day and on all Calendar Days of the year" at rank 2. "Can I send an instant transfer on a Sunday or on a bank holiday?" finds it at rank 46. The passage never says Sunday, weekend or holiday. It says "Calendar Days". A reader knows those are the same thing; the embedder puts them 44 places apart.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dropping duplicates, on the full set
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phrasing&lt;/th&gt;
&lt;th&gt;vector&lt;/th&gt;
&lt;th&gt;duplicates dropped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;original&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plain&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;terse&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One question flips, and it is Mikhail's, from rank 7 to 4 under both his first wording and the terse one. No other rank changes on the 21 originals, and every other move on the plain and terse phrasings is one to three places, all of them outside the top five already. The duplicate annex is 28.5% of the index, but it only crowds the top five when the question is about the annex, and none of the 20 golden questions is. H2 was wrong: dropping duplicates is the right fix for exactly one question in this set.&lt;/p&gt;

&lt;h3&gt;
  
  
  The hybrid path and the rule I promised
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Leg&lt;/th&gt;
&lt;th&gt;original&lt;/th&gt;
&lt;th&gt;plain&lt;/th&gt;
&lt;th&gt;terse&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;vector&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hybrid (production shape)&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hybrid, passage must match 2 words&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hybrid, passage must match 3 words&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the original questions the hybrid path loses three: Recall deadlines (rank 4 to 7), value limits (4 to 10), Return deadline (2 to 7). They are the same three the language piece lost, so that measurement reproduces. On the plain questions it loses six more, because "bank", "money" and "transfer" match passages all over both rulebooks. On the terse questions it gains two, since a three-word query is a lexical job and the full-text leg does it well.&lt;/p&gt;

&lt;p&gt;The minimum-words rule does nothing to the three lost questions. The rows are identical. I had expected the noise to be passages that share one acronym with the question, and a threshold of two would have removed them. That is not what the full-text leg returns for an English question:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["Does the SCT scheme itself&amp;lt;br/&amp;gt;set a maximum transaction&amp;lt;br/&amp;gt;amount? (6 query words)"] --&amp;gt; FT["full-text top 20"]
    FT --&amp;gt; N1["rank 1: SCT Inst p78&amp;lt;br/&amp;gt;obligations list&amp;lt;br/&amp;gt;matches 6 of 6 words"]
    FT --&amp;gt; N2["ranks 2 to 10&amp;lt;br/&amp;gt;match 5 of 6 words"]
    FT --&amp;gt; N3["ranks 11 to 20&amp;lt;br/&amp;gt;match 4 of 6 words"]
    FT -.-&amp;gt; G["labelled page, SCT p21&amp;lt;br/&amp;gt;Value Limits&amp;lt;br/&amp;gt;rank 93, matches 3 of 6"]
    style G fill:#FBEAE3,stroke:#993C1D&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Every passage in the full-text top 20 matches four to six of the six query words. The labelled page matches three and sits at rank 93. The same shape holds for the other two lost questions: the top 20 match four or five words each, the answer matches four and sits at rank 16 and 32. A threshold of two or three removes nothing, and the fusion goes on trusting a list whose top is wrong.&lt;/p&gt;

&lt;p&gt;So the loss on English questions is not single-word noise. It is the lexical leg ranking the right passage poorly, because a long question shares many ordinary words ("scheme", "set", "transaction", "amount") with many passages, and the one passage that answers it is not the one that repeats the most of them. The acronym explanation from the language piece was right for foreign-language questions, where the acronyms were the only words that could match, and wrong for English ones. H3 fails, and the rule I promised in the comments does not touch the problem it was meant to fix.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mikhail's question, all legs
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phrasing&lt;/th&gt;
&lt;th&gt;vector&lt;/th&gt;
&lt;th&gt;duplicates dropped&lt;/th&gt;
&lt;th&gt;full-text alone&lt;/th&gt;
&lt;th&gt;hybrid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;his rewording (original)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;his first question (plain)&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;terse&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This row reproduces the previous piece to the digit. It is also the only row in the set where the full-text leg has the answer at rank 1 under every phrasing: his first question has 9 words and the passage that holds the rule contains 8 of them, "PSMB", "telephone" and "meeting" among them. The lexical leg rescued his question because his question is the kind the lexical leg is good at. It does not follow that it rescues the others, and on this set it does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Wording moves retrieval more than any model swap did. A plain rewording of the same need costs 3 answers in 21; a search-box fragment costs 8 net, 9 lost and 1 gained. The model matrix piece found 11 of 80 questions depending on the model choice; here 10 of 21 depend on how the question is put, with the embedder held fixed.&lt;/li&gt;
&lt;li&gt;The loss is not "rulebook words good, plain words bad". Plain wording improved three ranks and lost three hits. What decides it is whether the words the user chose happen to appear in the passage: "Sunday" against "Calendar Days" loses, "twice by mistake" against "Duplicate sending" wins.&lt;/li&gt;
&lt;li&gt;Dropping exact duplicates is a fix for one question in this set, the one that is about the duplicated annex. It costs nothing and it is still worth doing, but it is not the reason the golden set misses.&lt;/li&gt;
&lt;li&gt;The hybrid path helps short queries and hurts long ones, and the minimum-matched-words rule does not change that, because on English questions the full-text leg ranks the right passage poorly rather than letting noise in. That corrects my own explanation from the language piece.&lt;/li&gt;
&lt;li&gt;The test type Mikhail asked for is worth keeping: same question, no section words, scored by the rank of the labelled page. It caught the weekend question, which passed every previous test in this series because every previous test used the rulebook's own words.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Whether the model answers correctly from the plain phrasing when the passage is retrieved. This piece stops at the ranking.&lt;/li&gt;
&lt;li&gt;The language-match rule (H4), which needs the Spanish and Russian questions from the language piece. Not run.&lt;/li&gt;
&lt;li&gt;A better lexical ranker (BM25 instead of &lt;code&gt;ts_rank&lt;/code&gt;, or field weights) as the fix the hybrid path actually needs. Argued from three questions, not built.&lt;/li&gt;
&lt;li&gt;More phrasings per question. Three is enough to show the effect and too few to give a per-question stability number.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The runner and the phrasing set, in the payments-rag repo: &lt;code&gt;comparison/paraphrase/run.py&lt;/code&gt;, &lt;code&gt;comparison/paraphrase/golden/sepa-phrasings.yaml&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Page labels: &lt;code&gt;evals/retrieval_golden_set.yaml&lt;/code&gt;, verified against the PDFs.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/the-answer-was-at-rank-7-2o81"&gt;The answer was at rank 7&lt;/a&gt; (the refused question, the duplicate annex).&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/asking-a-rag-in-the-wrong-language-885"&gt;Asking a RAG in the wrong language&lt;/a&gt; (hybrid at a cost of 2 to 3 in 20, the acronym explanation this piece corrects).&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l"&gt;Swapping every model in a RAG&lt;/a&gt; (11 of 80 questions model-dependent).&lt;/li&gt;
&lt;li&gt;EPC SCT and SCT Inst Rulebooks 2025 (the corpus).&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
      <category>python</category>
    </item>
    <item>
      <title>The answer was at rank 7</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:00:00 +0000</pubDate>
      <link>https://dev.to/skucherenko/the-answer-was-at-rank-7-2o81</link>
      <guid>https://dev.to/skucherenko/the-answer-was-at-rank-7-2o81</guid>
      <description>&lt;p&gt;A RAG system answers a question in two steps: a retriever picks a handful of passages out of the documents, and a language model writes an answer from those passages. When the answer is wrong, one of the two failed, and most evaluations, mine included, only score the answer. This piece adds the missing measurement, whether the passage that holds the answer was among the passages the model received, and re-reads an earlier experiment with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The earlier experiment
&lt;/h2&gt;

&lt;p&gt;The RAG in question indexes the two 2025 SEPA credit transfer rulebooks (the rules European banks follow for euro transfers, published by the EPC) in &lt;code&gt;Postgres&lt;/code&gt; with &lt;code&gt;pgvector&lt;/code&gt;. The text is split into 300-word chunks with a 50-word overlap, 484 chunks in total, and a question is answered from the 5 nearest chunks by cosine similarity. There is a &lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l"&gt;the model matrix article&lt;/a&gt; I held that pipeline fixed and swapped only the models. Six pipelines, each one an embedder (the model that turns text into vectors for the search) paired with a generator (the model that writes the answer). Four pipelines share &lt;code&gt;text-embedding-3-small&lt;/code&gt; as embedder, one uses &lt;code&gt;voyage-4&lt;/code&gt;, one uses &lt;code&gt;gemini-embedding-001&lt;/code&gt;; the generators are &lt;code&gt;Claude Sonnet&lt;/code&gt;, &lt;code&gt;Claude Haiku&lt;/code&gt;, &lt;code&gt;GPT&lt;/code&gt;, &lt;code&gt;Gemini Flash&lt;/code&gt; and &lt;code&gt;DeepSeek&lt;/code&gt;. Each pipeline answered 80 questions across four small corpora (SEPA rules, software licenses, agent-building guides, WHO nutrition sheets), each question with a reference answer written by hand, and a separate LLM judge scored every answer against its reference from 0 to 100, pass at 70. &lt;strong&gt;The result: 66 questions passed on all six pipelines, 3 failed on all six, and only 11 depended on which models were used&lt;/strong&gt;. I summarised it as "the corpus decides more than the model does".&lt;/p&gt;

&lt;p&gt;Three readers pushed on that. Edward Izgorodin asked to split those 3 + 11 questions by whether the answer passage reached the top 5 at all, because a judge score cannot tell a retrieval miss from a generator miss, and "whether the passage was in the top 5 is a lookup, not a verdict". Mikhail ran the live demo by hand, got a refusal with zero citations on a question whose answer is in the rulebook, saw two rewordings answer it correctly, and could not tell from the UI which component had failed. kaziava described a negative test that passed after an embedder swap because the trap chunk stopped being retrieved and the model never saw the bait.&lt;/p&gt;

&lt;p&gt;All three want the same number at the same point in the pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["question&amp;lt;br/&amp;gt;(80 in the matrix, each with&amp;lt;br/&amp;gt;a reference answer)"] --&amp;gt; R["retriever&amp;lt;br/&amp;gt;embedder + cosine search&amp;lt;br/&amp;gt;(3 embedders)"]
    R --&amp;gt; T["top 5 passages"]
    T --&amp;gt; G["generator&amp;lt;br/&amp;gt;(6 pipelines = embedder x LLM)"] --&amp;gt; A["answer"]
    A --&amp;gt; J["judge LLM&amp;lt;br/&amp;gt;answer vs reference&amp;lt;br/&amp;gt;score 0-100, pass at 70"]
    T -. "new measurement:&amp;lt;br/&amp;gt;answer phrase in the top 5? yes / no" .-&amp;gt; L["lookup&amp;lt;br/&amp;gt;(regex, no model)"]
    J --&amp;gt; V["verdict per pipeline"]
    L --&amp;gt; V&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The judge is a model reading a model; the lookup is a string search. Read together, they say which component failed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: a few terms from the rulebooks appear below. SCT is the standard SEPA Credit Transfer scheme; SCT Inst is the instant version; the PSMB is the EPC's Payment Scheme Management Board, whose meeting rules are an annex in both rulebooks. None of the findings depend on knowing more than that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Most of the 14 non-unanimous matrix questions would turn out to be retrieval misses. "The corpus decides" reads most naturally as "the passage was not there".&lt;/li&gt;
&lt;li&gt;Mikhail's refused question would be a generator failure: the passage in the top 5, the model too strict. Two rewordings answered it, and rewordings should not move a 484-chunk retriever much.&lt;/li&gt;
&lt;li&gt;Duplicated text between the two rulebooks would be a curiosity, a few pages of boilerplate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first held for about half the questions, the second turned out wrong, and the third was wrong by a wider margin than I would have guessed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The lookup
&lt;/h3&gt;

&lt;p&gt;The matrix has six pipelines but only three retrievers, since retrieval depends on the embedder alone: the four pipelines on &lt;code&gt;text-embedding-3-small&lt;/code&gt; receive the same five passages for a given question. So the 14 non-unanimous questions are 14 x 3 = 42 retrieval cells, and each cell either contains the answer passage or does not.&lt;/p&gt;

&lt;p&gt;"Contains" is a regex: a phrase from the reference answer that the source text must hold, one per question, checked against the five retrieved passages. No model. Two anchors were wrong on the first pass and I corrected them by reading the passages: the charging rule is written as "shared principle" in the rulebook, not &lt;code&gt;SHARE&lt;/code&gt; as in my reference answer, and "patent license" matched the GPL's patent clause when I was looking for Apache's. The anchor list is in the repo with the script.&lt;/p&gt;

&lt;p&gt;The judge scores are the ones from the matrix, &lt;code&gt;gpt&lt;/code&gt; as judge, pass bar 70.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mikhail's runs
&lt;/h3&gt;

&lt;p&gt;I read the stored production vectors and ran his exact questions through the same cosine search the live demo uses, k = 5, plus two variants: the same ranking with byte-identical chunks collapsed to one, and the hybrid path (vector plus &lt;code&gt;Postgres&lt;/code&gt; full-text search, &lt;code&gt;english&lt;/code&gt; configuration, the two rankings fused by reciprocal rank fusion). For each query I recorded the rank of the chunk that holds the answer, over all 484 chunks and not just the top 5.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: everything here reads the production database through a read-only connection and writes nothing to it. The live endpoint is plain vector retrieval, k = 5; dedup and hybrid are what it would do, not what it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The 14 questions, by retrieval
&lt;/h3&gt;

&lt;p&gt;Retrieved = rank of the first passage in the top 5 that holds the anchor phrase, per embedder. Scores are the judge's, pipelines p1 to p6.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;question&lt;/th&gt;
&lt;th&gt;oai / voyage / gem&lt;/th&gt;
&lt;th&gt;scores p1..p6&lt;/th&gt;
&lt;th&gt;what decided it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;lic-bsd3-third-clause&lt;/td&gt;
&lt;td&gt;none / none / none&lt;/td&gt;
&lt;td&gt;0,0,0,0,0,0&lt;/td&gt;
&lt;td&gt;retrieval, every embedder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sct-return-deadline&lt;/td&gt;
&lt;td&gt;none / none / none&lt;/td&gt;
&lt;td&gt;0,0,0,0,0,0&lt;/td&gt;
&lt;td&gt;retrieval, every embedder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lic-mpl-larger-work&lt;/td&gt;
&lt;td&gt;none / 2 / 4&lt;/td&gt;
&lt;td&gt;0,0,0,0,0,0&lt;/td&gt;
&lt;td&gt;half retrieval: the definition was there, the permission clause was not&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ag-langgraph-send&lt;/td&gt;
&lt;td&gt;none / 1 / 1&lt;/td&gt;
&lt;td&gt;0,0,0,100,100,0&lt;/td&gt;
&lt;td&gt;embedder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lic-apache-patent-retaliation&lt;/td&gt;
&lt;td&gt;none / 3 / 4&lt;/td&gt;
&lt;td&gt;0,0,0,100,100,0&lt;/td&gt;
&lt;td&gt;embedder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;nut-saturated-fat-limit&lt;/td&gt;
&lt;td&gt;none / none / 1&lt;/td&gt;
&lt;td&gt;0,0,0,0,100,0&lt;/td&gt;
&lt;td&gt;embedder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sct-max-execution-time&lt;/td&gt;
&lt;td&gt;none / none / 5&lt;/td&gt;
&lt;td&gt;0,0,0,0,100,0&lt;/td&gt;
&lt;td&gt;embedder, plus a chunk cut before the number&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lic-apache-patent&lt;/td&gt;
&lt;td&gt;none / none / none&lt;/td&gt;
&lt;td&gt;0,0,100,0,0,65&lt;/td&gt;
&lt;td&gt;retrieval miss on every embedder, and one pass anyway&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ag-aci&lt;/td&gt;
&lt;td&gt;1 / 1 / 1&lt;/td&gt;
&lt;td&gt;100,100,93,100,100,55&lt;/td&gt;
&lt;td&gt;generator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ag-openai-guardrails&lt;/td&gt;
&lt;td&gt;1 / 1 / 2&lt;/td&gt;
&lt;td&gt;86,71,50,71,100,72&lt;/td&gt;
&lt;td&gt;generator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ag-openai-when-agent&lt;/td&gt;
&lt;td&gt;4 / 3 / 1&lt;/td&gt;
&lt;td&gt;98,65,95,100,100,0&lt;/td&gt;
&lt;td&gt;generator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lic-gpl-installation-info&lt;/td&gt;
&lt;td&gt;1 / 1 / 1&lt;/td&gt;
&lt;td&gt;70,65,78,65,100,45&lt;/td&gt;
&lt;td&gt;generator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sct-charging-principle&lt;/td&gt;
&lt;td&gt;1 / 1 / 2&lt;/td&gt;
&lt;td&gt;100,100,0,95,100,100&lt;/td&gt;
&lt;td&gt;generator, confused by a duplicate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sct-inquiry-reasons&lt;/td&gt;
&lt;td&gt;2 / 2 / 1&lt;/td&gt;
&lt;td&gt;67,65,70,67,75,67&lt;/td&gt;
&lt;td&gt;judge threshold&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Six of the fourteen are decided by retrieval. For four of them the scores follow the embedder exactly: the passage is absent on the embedders that scored 0 and present on the ones that scored 100, with nothing in between. &lt;code&gt;LangGraph&lt;/code&gt;'s Send API and the Apache retaliation clause are found by &lt;code&gt;voyage-4&lt;/code&gt; and &lt;code&gt;gemini-embedding-001&lt;/code&gt; and not by the production embedder; the WHO saturated fat limit and the SCT execution time only by Gemini. Two more were never retrieved by anyone, so no generator could have answered them, and one of those is the BSD-3 question, where the license name exists only in the filename, which no embedder sees.&lt;/p&gt;

&lt;p&gt;Five are the generator's, with the passage at rank 1 or 2 on every embedder. &lt;code&gt;DeepSeek&lt;/code&gt; refused the OpenAI workflow question with the passage at rank 4 and the judge said so: "incorrectly declines to answer, since the context explicitly identifies workflows involving complex decisions, unstructured data, and brittle rule-based systems". The inquiry-reasons question has all six scores within 5 points of the pass bar, which is not a model difference, it is where I drew the line.&lt;/p&gt;

&lt;p&gt;The Apache patent grant is the row Edward predicted and kaziava had already lived through. The grant clause was not in the top 5 on any embedder. Four generators refused, correctly. &lt;code&gt;gpt&lt;/code&gt; answered "Yes. The Apache-2.0 license includes a patent grant, though the provided excerpt does not contain its specific terms." and the judge gave it 100, flagged grounded=false, because the yes/no was right. That is a pass for the wrong reason, a fact retrieved from the model's memory rather than the corpus, and the pass rate hides it completely where the lookup shows it in one column.&lt;/p&gt;

&lt;p&gt;The charging-principle row is the one that sent me to the duplicate count. &lt;code&gt;gpt&lt;/code&gt; scored 0 with the passage at rank 1 because it read "The available charging principle applies to SCT Inst only". The chunk it received was the SCT Inst rulebook's copy of the charging section; the SCT rulebook has the same paragraph, word for word, and it was not in the top 5 because its twin had taken the slot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mikhail's question
&lt;/h3&gt;

&lt;p&gt;His original question, "Within how many calendar days can a PSMB member request a telephone meeting?", was refused twice on the live demo with zero citations. Against the stored vectors the passage that holds the rule is at rank 7. The top 5 were the general PSMB meeting pages, and 4 of those 5 slots were two duplicate pairs, so the generator saw three distinct passages and none of them had the number. A retrieval miss that the prompt turned into a refusal, exactly as he guessed from the outside.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["'Within how many calendar days can a&amp;lt;br/&amp;gt;PSMB member request a telephone meeting?'"] --&amp;gt; R["vector search, k = 5"]
    R --&amp;gt; T["top 5: general PSMB meeting pages&amp;lt;br/&amp;gt;(two duplicate pairs + one)&amp;lt;br/&amp;gt;rule passage at rank 7"]
    T --&amp;gt; G["generator: no number in context"] --&amp;gt; A["'the sources do not contain information'"] --&amp;gt; U["UI: Evidence 0 passages&amp;lt;br/&amp;gt;(retrieved passages not shown)"]
    R -. "collapse duplicates" .-&amp;gt; D["rule passage at rank 4"]
    R -. "add full-text leg, fuse" .-&amp;gt; H["rule passage at rank 3"]
    R -. "reword with 'written procedure'" .-&amp;gt; W["rule passage at rank 1"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;His two rewordings, which add "instead of the written vote" or "after receiving a written communication with a proposed decision", put the same passage at rank 1. The words that carry the section are "written procedure"; his original question does not contain them, and the rulebook does not contain "telephone meeting" as a phrase either (it says "PSMB meeting by telephone"). The rule is the last sentence of a 300-word chunk about written votes, so the chunk as a whole is about written votes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;query&lt;/th&gt;
&lt;th&gt;vector, rank of the rule&lt;/th&gt;
&lt;th&gt;duplicates collapsed&lt;/th&gt;
&lt;th&gt;hybrid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;original (telephone meeting, calendar days)&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"... instead of the written vote?"&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"... after receiving a written communication with a proposed decision?"&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Voting by written procedure"&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"The communication shall be" (a 4-word quote)&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Either of the two fixes he proposed would have answered his original question on its own. Collapsing exact duplicates before building the context moves the rule from rank 7 to rank 4. The full-text leg alone had it at rank 1, and fusing it with the vector list puts it at rank 3. The 4-word quote is the case full-text exists for: vector rank 25, dedup leaves it there, full-text rank 1.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: in &lt;a href="https://dev.to/skucherenko/asking-a-rag-in-the-wrong-language-885"&gt;the previous piece&lt;/a&gt; the same hybrid path cost 2 to 3 questions in 20 on my golden set, in every language, because the lexical leg matched acronyms all over the corpus and the fusion trusted it. That measurement stands. Mikhail's questions are of a different kind, terse and section-specific, and there the lexical leg is the better retriever. The honest position is that the choice depends on the query, which is what he wrote in his comment about his own code-search ablation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The duplicates
&lt;/h3&gt;

&lt;p&gt;484 chunks in the index. 138 of them, 28.5%, are byte-identical to a chunk from the other rulebook. Both 2025 rulebooks bind the same 36-page "EPC Payment Scheme Management Rules" document as an annex, and my ingestion, which works one PDF at a time, indexed it twice. Every one of the 69 duplicate groups is one copy per rulebook.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["sct_rulebook_2025.pdf&amp;lt;br/&amp;gt;(248 chunks)"] --&amp;gt;|"from page 93:&amp;lt;br/&amp;gt;Scheme Management Rules, 36 pages"| X["chunker + embedder,&amp;lt;br/&amp;gt;one PDF at a time"]
    B["sct_inst_rulebook_2025.pdf&amp;lt;br/&amp;gt;(236 chunks)"] --&amp;gt;|"from page 95:&amp;lt;br/&amp;gt;the same 36 pages"| X
    X --&amp;gt; I["index: 484 chunks,&amp;lt;br/&amp;gt;138 of them in identical pairs (28.5%)"]
    I --&amp;gt; T["a top 5 on a PSMB question:&amp;lt;br/&amp;gt;pair, pair, one single&amp;lt;br/&amp;gt;= 3 distinct passages"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This is why the duplicate pair took 2 of 5 slots in Mikhail's "voting by written procedure" query (in my reproduction, 4 of 5, two pairs), why &lt;code&gt;gpt&lt;/code&gt; thought the charging rule was SCT Inst only, and why the previous piece's page labels undercounted one question: the label for the charging question names the SCT rulebook page, the retriever returns the identical SCT Inst page at rank 1 for all three embedders, and the eval scored it a miss. By passage the production embedder misses 3 of 20 English questions, not 4.&lt;/p&gt;

&lt;p&gt;The fix Mikhail proposed is the right one: collapse identical text before building the context and keep both sources in the citation, since "this rule applies to both schemes" is information a payments engineer wants.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Whether the passage was retrieved is a lookup, and it changes how the matrix reads. Of the 14 questions that were not unanimous across six pipelines, 6 were decided by the embedder before any generator ran and a seventh got only half its passage, 5 by the generator with the passage in hand, 1 by where I put the pass bar, and 1 was a pass with the passage absent, which no judge score could have shown. The sentence "the corpus decides more than the model" survives, with a sharper meaning: it is mostly the retriever that decides, and the retriever is the embedder plus the chunker plus whatever the index contains twice.&lt;/p&gt;

&lt;p&gt;Mikhail's refusal was a retrieval miss at rank 7, dressed as a refusal by the prompt. The UI showed him only the passages the generator cited, so he had to guess. Showing retrieved passages next to cited ones costs nothing and would have made his comment a one-liner.&lt;/p&gt;

&lt;p&gt;A quarter of my index is duplicated, and I found out from a reader. Exact-duplicate collapsing goes in before the next measurement round.&lt;/p&gt;

&lt;p&gt;Retrieval traps deserve their own verdict. kaziava's "test did not run", for a negative test whose trap chunk was not retrieved, is the third state my eval was missing. The Apache patent row is what it looks like when that state is missing: a correct answer from memory, a judge that agrees, and no line in any report.&lt;/p&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Answer quality with duplicates collapsed, on the full golden set. This piece shows ranks for six queries; the recall@5 and judge numbers with dedup on are the next run.&lt;/li&gt;
&lt;li&gt;A query-type rule for when to fuse the lexical leg. Argued above from six queries, not measured.&lt;/li&gt;
&lt;li&gt;Paraphrase robustness across the 20 golden questions. Mikhail's three phrasings moved one passage from rank 7 to rank 1; how often that happens is its own measurement.&lt;/li&gt;
&lt;li&gt;The UI change. Recorded as a decision, not built.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The comments: Edward Izgorodin, Mikhail and kaziava under &lt;a href="https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l"&gt;Swapping every model in a RAG&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;kaziava, &lt;a href="https://dev.to/kaziava/the-negative-test-that-passed-for-the-wrong-reason-3koh"&gt;The Negative Test That Passed for the Wrong Reason&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l"&gt;Swapping every model in a RAG&lt;/a&gt; (the matrix: 80 questions, 6 pipelines, 3 judges)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/asking-a-rag-in-the-wrong-language-885"&gt;Asking a RAG in the wrong language&lt;/a&gt; (hybrid at a cost of 2 to 3 in 20; the page labels)&lt;/li&gt;
&lt;li&gt;GitHub repo: &lt;code&gt;comparison/matrix/&lt;/code&gt; (pipelines, golden sets, judge scores) and &lt;code&gt;comparison/multilingual/run.py&lt;/code&gt; (the read-only store used for the reproductions)&lt;/li&gt;
&lt;li&gt;EPC Payment Scheme Management Rules, EPC207-14 v5.0, annexed to both 2025 rulebooks&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
      <category>python</category>
    </item>
    <item>
      <title>Asking a RAG in the wrong language</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:11:20 +0000</pubDate>
      <link>https://dev.to/skucherenko/asking-a-rag-in-the-wrong-language-885</link>
      <guid>https://dev.to/skucherenko/asking-a-rag-in-the-wrong-language-885</guid>
      <description>&lt;p&gt;Every article in this series so far assumed that the person asking speaks the language of the documents. The SEPA rulebooks my &lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;RAG&lt;/a&gt; indexes exist only in English, and so did every question in my golden set. But a payments engineer in Madrid or Kyiv asks in their own language, and I had never measured what that does to retrieval. I can read Spanish and Russian myself, so those two were the cheap place to start.&lt;/p&gt;

&lt;p&gt;The question for this piece is narrow: what happens to the retriever when the question changes language and the corpus does not. It stops at whether the right page still lands in the top 5; answer quality and the generator wait for the next piece.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;p&gt;Three things I expected going in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vector retrieval would lose something cross-lingually, since the embedder has to map a Spanish question next to an English paragraph, but not much. Vendors describe these models as multilingual.&lt;/li&gt;
&lt;li&gt;Hybrid retrieval (vector plus Postgres full-text search, fused by reciprocal rank fusion) would quietly stop helping, because the full-text index is built with the &lt;code&gt;english&lt;/code&gt; configuration. My guess was that a Spanish query would match nothing and the fusion would fall back to the vector ranking alone.&lt;/li&gt;
&lt;li&gt;Translating the question to English first, then running the normal English pipeline, would recover most of the loss for the cost of one small LLM call.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second one turned out wrong in an instructive way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;

&lt;p&gt;The corpus is the production index untouched: 484 chunks from the two 2025 EPC rulebooks, with the contextual blurbs from &lt;a href="https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049"&gt;the previous improvement round&lt;/a&gt; already in the stored vectors. The eval reads it read-only.&lt;/p&gt;

&lt;p&gt;The golden set is the 20 answerable SEPA questions from the model matrix article, each now labelled with the page that answers it, and each translated into Spanish and Russian. Scheme and message names (&lt;code&gt;SCT&lt;/code&gt;, &lt;code&gt;SCT Inst&lt;/code&gt;, &lt;code&gt;PSP&lt;/code&gt;, &lt;code&gt;Recall&lt;/code&gt;, &lt;code&gt;Return&lt;/code&gt;) stay in English in all three versions, because that is how practitioners actually talk about them. It also turned out to matter for what follows.&lt;/p&gt;

&lt;p&gt;Every question, in each of its three languages, goes through four retrieval paths, and each path hands back a top 5 that is scored against the page labels:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Vector only: embed the question, take the 5 nearest chunks. Run once per embedder.&lt;/li&gt;
&lt;li&gt;Full-text only: Postgres full-text search, the 5 best lexical matches. Run once per language configuration.&lt;/li&gt;
&lt;li&gt;Hybrid: take the top 20 from vector and the top 20 from full-text, fuse the two rankings, keep the top 5. This is the production hybrid path with one change: the language configuration is a variable instead of the hardcoded &lt;code&gt;english&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Translate first: &lt;code&gt;Claude Haiku&lt;/code&gt; turns the Spanish or Russian question into English, then path 1 runs with the production embedder.
&lt;/li&gt;
&lt;/ol&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["question&amp;lt;br/&amp;gt;(en / es / ru)"] --&amp;gt; V["1. vector only"] --&amp;gt; T5a["top 5"]
    Q --&amp;gt; F["2. full-text only"] --&amp;gt; T5b["top 5"]
    Q --&amp;gt; H["3. hybrid: vector top-20 + full-text top-20,&amp;lt;br/&amp;gt;fused by reciprocal rank"] --&amp;gt; T5c["top 5"]
    Q --&amp;gt; TR["4. translate to English,&amp;lt;br/&amp;gt;then vector only"] --&amp;gt; T5d["top 5"]
    T5a &amp;amp; T5b &amp;amp; T5c &amp;amp; T5d --&amp;gt; S["scored against&amp;lt;br/&amp;gt;the page labels"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The fusion in path 3 is reciprocal rank fusion, the same one production uses. Each chunk scores 1 / (60 + its rank) in every list it appears in, and the two scores are added. A chunk at rank 1 in one list and absent from the other gets 1/61; a chunk at rank 10 in both gets 2/70, which is more. So the fusion rewards agreement between the two lists and treats both lists as equally trustworthy. That second property is the one this piece is about.&lt;/p&gt;

&lt;p&gt;The three embedders are the ones from the matrix piece: &lt;code&gt;text-embedding-3-small&lt;/code&gt; (production, its vectors read straight from the database), &lt;code&gt;voyage-4&lt;/code&gt; and &lt;code&gt;gemini-embedding-001&lt;/code&gt;. The full-text configurations tried were &lt;code&gt;english&lt;/code&gt;, &lt;code&gt;spanish&lt;/code&gt;, &lt;code&gt;russian&lt;/code&gt; and &lt;code&gt;simple&lt;/code&gt; (no stemming at all).&lt;/p&gt;

&lt;p&gt;Two metrics. hit@5 against the page labels is the real one; with 20 questions every hit is worth 0.05. The second needs no labels at all: how much of the top 5 for the Spanish or Russian question overlaps with the top 5 for the same question in English. It says whether the retriever is even looking at the same pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: the production &lt;code&gt;/ask&lt;/code&gt; endpoint runs plain vector retrieval. Hybrid exists behind an eval flag, since the round where I measured it landed at &lt;a href="https://www.linkedin.com/pulse/illusion-improvement-serhiy-kucherenko-vw20e/" rel="noopener noreferrer"&gt;exactly the same recall as vector alone&lt;/a&gt;. So nothing below describes a live failure; it describes what turning that flag on would do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;hit@5 over 20 questions, production embedder unless stated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;leg&lt;/th&gt;
&lt;th&gt;en&lt;/th&gt;
&lt;th&gt;es&lt;/th&gt;
&lt;th&gt;ru&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;vector, &lt;code&gt;text-embedding-3-small&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.65&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vector, &lt;code&gt;gemini-embedding-001&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;0.85&lt;/td&gt;
&lt;td&gt;0.75&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vector, &lt;code&gt;voyage-4&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;full-text alone, &lt;code&gt;english&lt;/code&gt; config&lt;/td&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;td&gt;0.35&lt;/td&gt;
&lt;td&gt;0.35&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hybrid (vector + full-text), &lt;code&gt;english&lt;/code&gt; config&lt;/td&gt;
&lt;td&gt;0.65&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;0.55&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hybrid, &lt;code&gt;spanish&lt;/code&gt; config&lt;/td&gt;
&lt;td&gt;0.65&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;translate to English, then vector&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.75&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Top-5 overlap with the English question: 0.72 for Spanish, 0.63 for Russian on the production embedder; 0.81 and 0.73 on the &lt;code&gt;Gemini&lt;/code&gt; embedder; 0.82 and 0.76 on &lt;code&gt;Voyage&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The cross-lingual loss is real and small. Spanish costs the production embedder 2 questions in 20, Russian 3. English itself misses 4 of the 20; Spanish keeps all four of those misses and adds two, Russian keeps three of them and adds four, so the language penalty sits mostly on top of what English already gets wrong. &lt;code&gt;Gemini&lt;/code&gt; also loses 2 and 3; &lt;code&gt;Voyage&lt;/code&gt; starts lower in English and loses 0 and 2. Meaning, the embedder you already have is not the problem here, or at least not the biggest one.&lt;/p&gt;

&lt;p&gt;Hybrid costs recall in every language, English included. Against plain vector it drops 0.80 to 0.65 in English, 0.70 to 0.60 in Spanish, 0.65 to 0.55 in Russian. In English it loses 3 hits and gains none; in Spanish it loses 2 and gains none; in Russian it loses 3 and gains 1. The English drop is new since the round where hybrid was at parity: contextual retrieval lifted the vector leg from 0.60 to 0.80 in between, and the full-text leg stayed where it was, so fusing them now dilutes a better ranking with a worse one.&lt;/p&gt;

&lt;p&gt;The mechanism is where my second hypothesis broke. I expected a Spanish question to match nothing in an English full-text index. It matched 192 of 484 chunks. The production query rewrites the full-text search from AND to OR (a natural-language question shares only some words with a terse rulebook paragraph, so requiring every word matched nothing), and under OR any shared token is enough. For the question about Recall deadlines the &lt;code&gt;english&lt;/code&gt; stemmer turned the Spanish text into this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;'en' | 'qué' | 'plazo' | 'debe' | 'enviars' | 'un' | 'recal' | 'de' | 'sct' | 'por' | 'duplicidad' | 'o' | 'error' | 'técnico' | 'frent' | 'fraud'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;recal&lt;/code&gt;, &lt;code&gt;sct&lt;/code&gt;, &lt;code&gt;error&lt;/code&gt; and &lt;code&gt;fraud&lt;/code&gt; survive translation, so 192 chunks match, ranked by how often they mention those tokens. The full-text leg returned all 20 of its 20 candidates on every one of the Spanish questions, and on 19 of the 20 Russian ones (one Russian question shared no token at all and matched zero, which is the behaviour I had expected everywhere). The top of that list for the Recall question was the SCT Inst page listing Recall reasons, the change log, and an annex, all of them about Recalls and none of them the page with the deadlines. Reciprocal rank fusion then weights that list equally with the vector list, so twenty plausible but wrong pages get the same say as the twenty the embedder picked.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q["Spanish question&amp;lt;br/&amp;gt;¿En qué plazos debe&amp;lt;br/&amp;gt;enviarse un Recall...?"] --&amp;gt; TS["plainto_tsquery('english'),&amp;lt;br/&amp;gt;AND rewritten to OR&amp;lt;br/&amp;gt;recal | sct | error | fraud | plazo | ..."]
    TS --&amp;gt; FT["192 of 484 chunks match&amp;lt;br/&amp;gt;full-text top-20 by ts_rank:&amp;lt;br/&amp;gt;SCT Inst Recall p.40, change log p.8,&amp;lt;br/&amp;gt;annex p.134, ..."]
    Q --&amp;gt; VEC["vector top-20&amp;lt;br/&amp;gt;deadline page p.31 at rank 3"]
    FT --&amp;gt; RRF["RRF fusion&amp;lt;br/&amp;gt;both lists weighted equally"]
    VEC --&amp;gt; RRF
    RRF --&amp;gt; OUT["hybrid top-5&amp;lt;br/&amp;gt;five SCT Inst Recall pages,&amp;lt;br/&amp;gt;p.31 pushed out"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;For that question the vector leg alone had the deadline page at rank 3. After fusion the top 5 were five pages from the SCT Inst Recall section, and the deadline page was gone.&lt;/p&gt;

&lt;p&gt;Switching the full-text configuration to &lt;code&gt;spanish&lt;/code&gt; or &lt;code&gt;russian&lt;/code&gt; changed nothing for Spanish or Russian questions (0.60 and 0.55 to 0.60, within one question). The documents are English, so a Spanish stemmer has nothing to stem on the index side; the only cross-language matches were the acronyms, and those match under any configuration. The fix is not in the stemmer.&lt;/p&gt;

&lt;p&gt;Translating first helped Russian and did nothing for Spanish, which looked odd, since Spanish is the closer language. &lt;code&gt;Claude Haiku&lt;/code&gt; turned the 40 questions into English for $0.0071 total, about $0.0002 and 0.7 seconds per query. Russian went from 0.65 to 0.75; Spanish stayed at 0.70 with exactly the same hits. Top-5 overlap with the English question rose to 0.87 and 0.81 for both, so translation did put the retriever back on the same pages even where the count did not move.&lt;/p&gt;

&lt;p&gt;The explanation is in the ranks, not in the languages. The two questions that separate the Spanish and Russian translate legs are ones where the original English question itself only just makes it: the answer page sits at rank 4 for one and rank 5 for the other. A translation is a paraphrase, and any paraphrase can move a rank-5 page to rank 6. On the value-limits question the Russian back-translation came out word for word identical to my English original (rank 4, hit); the Spanish one came out as "establish a maximum amount per transaction" (rank 6, miss). On the execution-time question the Spanish back-translation dropped the word "credit" from "credit transfer", and the page went from rank 5 in English to rank 13. Russian also had more to recover: native Spanish already found two pages (settlement certainty, currency) at ranks 1 and 3 that native Russian had at rank 7. Meaning, translation buys back what the language cost, and then the fragile English questions cost it right back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: I also wanted to say there were no false positives, and I cannot. The golden set carries 5 questions whose answer is verifiably not in the corpus (they are about direct debits, card fees, SWIFT messages and TARGET2 hours). I ran them in all three languages through the production answer path: 14 of 15 were refused correctly, in the language of the question. The 15th was the Russian version of the direct-debit refund question, which came back with "13 months" and a citation. Thirteen months is the window for recalling a SEPA Credit Transfer, a different payment instrument entirely; the English and Spanish versions of the same question were refused. One hallucination in fifteen, and it appeared only cross-lingually.&lt;/p&gt;

&lt;p&gt;Where the same knob lives on other stacks, for anyone running one of the systems from &lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-1-framework-serhiy-kucherenko-tbdoe/" rel="noopener noreferrer"&gt;the comparison&lt;/a&gt; with lexical search turned on (the comparison itself ran them vector-only):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Postgres&lt;/code&gt;: the text search configuration is the second argument of &lt;code&gt;to_tsvector&lt;/code&gt; and &lt;code&gt;plainto_tsquery&lt;/code&gt;; the index is built with one of them too, so a per-query change means either a second index or a sequential scan.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LlamaIndex&lt;/code&gt;: &lt;code&gt;BM25Retriever&lt;/code&gt; takes &lt;code&gt;stemmer=Stemmer.Stemmer("english")&lt;/code&gt; and &lt;code&gt;language="english"&lt;/code&gt; for the stop-word list, both defaulting to English.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Haystack&lt;/code&gt;: the in-memory BM25 lowercases and tokenizes with the regex &lt;code&gt;(?u)\b\w+\b&lt;/code&gt;, and that is all. No stemming, no stop words, no language to set, so a Spanish question matches an English corpus on exactly the shared tokens this piece is about.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LangChain&lt;/code&gt;: &lt;code&gt;BM25Retriever&lt;/code&gt; takes a &lt;code&gt;preprocess_func&lt;/code&gt;; the default splits on whitespace, and any stemming or stop-word handling is yours to write.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;OpenAI&lt;/code&gt; file search: the docs say retrieval is "semantic and keyword search" and expose no language, tokenizer or keyword setting. &lt;code&gt;NotebookLM&lt;/code&gt; exposes nothing at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In every one of these the lexical leg has a language baked in somewhere, and none of them will tell you when the question stops matching it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Cross-lingual retrieval with the production embedder loses 2 to 3 questions in 20 against English. Small, consistent across three embedders, and sitting on top of the 4 questions English misses on its own.&lt;/li&gt;
&lt;li&gt;Hybrid search is not neutral to language, and not for the reason I guessed. It does not fall back to vector when the query language changes; it fuses in a full list of confident lexical matches on shared acronyms, and with an OR-joined query that list is always full. It cost 2 to 3 questions in every language, English included.&lt;/li&gt;
&lt;li&gt;The full-text language configuration is the wrong knob when the documents and the questions disagree. The right knob is fusion: skip or down-weight the lexical leg when the query language is not the corpus language, or require more than one matching term before a chunk counts.&lt;/li&gt;
&lt;li&gt;Translating the question first is cheap ($0.0002, 0.7 s) and recovered Russian to within one question of English. Where it did not help, the reason was two English questions that already sit at rank 4 and 5, so a paraphrase pushes them out. Translation is a paraphrase.&lt;/li&gt;
&lt;li&gt;The refusal boundary that held for all 20 trap questions in the matrix piece let one through here, in Russian only. Cross-lingual questions deserve their own refusal traps in the golden set.&lt;/li&gt;
&lt;li&gt;Twenty questions means every hit is 0.05; I am reporting deltas of 2 and 3 questions and would not defend anything smaller. Honest rather than optimized.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The generator, beyond the 15 refusal checks above. Whether &lt;code&gt;Claude&lt;/code&gt; answers a Spanish question correctly from English passages, and in which language, is the next measurement.&lt;/li&gt;
&lt;li&gt;A corpus that is itself not in English. That is a different problem: the chunker and the embedder both see different text, and it is the subject of the next piece. The obvious shortcut, translating the whole corpus into one language before indexing, is a document-translation problem at scale, where every sentence has to stay correct and nobody signs off on it. The tooling that exists for that is quality estimation, models like &lt;a href="https://arxiv.org/abs/2209.06243" rel="noopener noreferrer"&gt;CometKiwi&lt;/a&gt; that score a translation without a reference, and even its authors &lt;a href="https://www2.statmt.org/wmt24/pdf/2024.wmt-1.121.pdf" rel="noopener noreferrer"&gt;warn about reading the scores as guarantees&lt;/a&gt;. A recent cross-lingual retrieval study (&lt;a href="https://arxiv.org/abs/2511.19324" rel="noopener noreferrer"&gt;Goworek et al., 2025&lt;/a&gt;) found that dense retrievers trained for cross-lingual use "derive little benefit from document translation" and recommends multilingual embeddings over translation pipelines. So the honest position for now: do not translate the corpus, measure the embedder on it.&lt;/li&gt;
&lt;li&gt;A fusion rule that is language-aware. Argued above, not implemented or measured.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;: &lt;code&gt;comparison/multilingual/&lt;/code&gt; (runner, golden sets in three languages, page labels)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.linkedin.com/pulse/illusion-improvement-serhiy-kucherenko-vw20e/" rel="noopener noreferrer"&gt;The illusion of improvement&lt;/a&gt; (hybrid at parity with vector, 0.60 = 0.60)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049"&gt;Improving the scores of a RAG&lt;/a&gt; (contextual retrieval, 0.60 to 0.80)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l"&gt;Swapping every model in a RAG&lt;/a&gt; (the three embedders and the 20 SEPA questions)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/textsearch-intro.html" rel="noopener noreferrer"&gt;Postgres full-text search: text search configurations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://developers.llamaindex.ai/python/framework/integrations/retrievers/bm25_retriever/" rel="noopener noreferrer"&gt;LlamaIndex BM25Retriever&lt;/a&gt; (&lt;code&gt;stemmer&lt;/code&gt;, &lt;code&gt;language&lt;/code&gt;), &lt;a href="https://docs.haystack.deepset.ai/docs/inmemorydocumentstore" rel="noopener noreferrer"&gt;Haystack InMemoryDocumentStore&lt;/a&gt; (&lt;code&gt;bm25_tokenization_regex&lt;/code&gt;), &lt;a href="https://reference.langchain.com/python/langchain-community/retrievers/bm25/BM25Retriever" rel="noopener noreferrer"&gt;LangChain BM25Retriever&lt;/a&gt; (&lt;code&gt;preprocess_func&lt;/code&gt;), &lt;a href="https://developers.openai.com/api/docs/guides/tools-file-search" rel="noopener noreferrer"&gt;OpenAI file search&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2511.19324" rel="noopener noreferrer"&gt;Goworek, Macmillan-Scott, Özyiğit: What Drives Cross-lingual Ranking?&lt;/a&gt; (2025)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2209.06243" rel="noopener noreferrer"&gt;CometKiwi&lt;/a&gt; and &lt;a href="https://www2.statmt.org/wmt24/pdf/2024.wmt-1.121.pdf" rel="noopener noreferrer"&gt;Pitfalls and Outlooks in Using COMET&lt;/a&gt; (translation quality estimation)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>rag</category>
      <category>database</category>
    </item>
    <item>
      <title>Swapping every model in a RAG</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Mon, 14 Sep 2026 19:08:51 +0000</pubDate>
      <link>https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l</link>
      <guid>https://dev.to/skucherenko/swapping-every-model-in-a-rag-5h8l</guid>
      <description>&lt;p&gt;In the comparison articles (&lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-1-framework-serhiy-kucherenko-tbdoe/" rel="noopener noreferrer"&gt;part 1&lt;/a&gt; and &lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-2-benchmark-serhiy-kucherenko-kflse/" rel="noopener noreferrer"&gt;part 2&lt;/a&gt;) I put my own &lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;RAG&lt;/a&gt; against five other systems. The models were frozen there on purpose: same embedder, same generator everywhere, so that only the framework could make a difference. The 'Out of Scope' section admitted the two gaps that the design left open: one golden set on one topic, and no model variation at all.&lt;/p&gt;

&lt;p&gt;This experiment inverts it. The pipeline is now the frozen part, a deliberately boring hand-rolled one, and the models are the only variables: three embedders, five generators, three judges, four topics. I am not looking for a winner here. If the behavior and the numbers survive swapping every model slot, then the conclusions belong to the system rather than to some vendor's checkpoint, and I can replace any model later without the evaluation story collapsing. And if they do not survive it, better to learn that now than after building more on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;p&gt;The first change is the ruler itself. The golden set grew from 10 questions on one topic to 100 across four topics. With 10 questions, one lucky or unlucky answer moved the score by 10 points; with 100, the numbers get stable enough that comparing models means something at all. 20 of the 100 are refusal traps: their answer is verifiably NOT in the corpus, so the only correct answer is "the documents do not cover this".&lt;/p&gt;

&lt;p&gt;Then, three questions going in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Do my numbers depend on the models I picked?&lt;/strong&gt; Everything so far was measured with one embedder and one generator. If swapping models changes the story, the numbers describe those models, not my system. And if models do differ, which slot matters more: the embedder or the generator?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Are the scores themselves model-independent?&lt;/strong&gt; Every answer is scored by a &lt;code&gt;Claude&lt;/code&gt;, a &lt;code&gt;GPT&lt;/code&gt; and a &lt;code&gt;Gemini&lt;/code&gt; judge, on purpose including the judge from the generator's own vendor. If a judge favors its own family, the scores are part of the problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does a RAG answer from the corpus or from the model's memory?&lt;/strong&gt; The refusal traps test exactly this: a model that answers them is answering from memory. Several are bait, sitting right next to things the corpus DOES mention.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;

&lt;p&gt;One frozen pipeline: 200-word chunks with 20-word overlap, cosine similarity, top-5 to the generator, one shared answer prompt that demands answering only from the provided context. No frameworks, no reranking, and the production database untouched (a lesson from the comparison piece that I keep respecting).&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q[question] --&amp;gt; E["embedder&amp;lt;br/&amp;gt;(slot: 3 models)"]
    E --&amp;gt; R["cosine top-5&amp;lt;br/&amp;gt;(frozen)"]
    R --&amp;gt; G["generator&amp;lt;br/&amp;gt;(slot: 5 models)"]
    G --&amp;gt; A[answer + sources]
    A --&amp;gt; J1["judge: Claude&amp;lt;br/&amp;gt;(scores all)"]
    A --&amp;gt; J2["judge: GPT&amp;lt;br/&amp;gt;(scores all)"]
    A --&amp;gt; J3["judge: Gemini&amp;lt;br/&amp;gt;(scores all)"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Four topics, chosen so that they are not four flavors of the same thing: the SEPA rulebooks from production (10 questions kept verbatim from the comparison golden set), open-source license texts plus the GNU GPL FAQ, two vendors' agent-building guides plus &lt;code&gt;LangGraph&lt;/code&gt; docs, and WHO nutrition fact sheets. Nutrition is there deliberately as the topic where a model's pretrained opinions are strongest, so refusing to invent an answer costs the most.&lt;/p&gt;

&lt;p&gt;Six pipelines; each one is an edge picking one embedder and one generator, with the same frozen retrieval between the two columns:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph EMB[embedders]
        E1["text-embedding-3-small&amp;lt;br/&amp;gt;(production)"]
        E2[Voyage-4]
        E3[Gemini embedding]
    end
    subgraph GEN[generators]
        G1[Claude Sonnet]
        G2[Claude Haiku]
        G3[GPT-5.6 Terra]
        G4[DeepSeek]
        G5[Gemini Flash]
    end
    E1 --&amp;gt;|"p1 (baseline)"| G1
    E1 --&amp;gt;|p2| G2
    E1 --&amp;gt;|p3| G3
    E1 --&amp;gt;|p6| G4
    E2 --&amp;gt;|p4| G1
    E3 -.-&amp;gt;|"p5 (whole stack)"| G5&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Reading it: p2, p3 and p6 keep p1's embedder and change only the generator; p4 keeps p1's generator and changes only the embedder. So every comparison against the baseline isolates exactly one component (a test enforces this). The one exception is p5, which changes both at once, so it is flagged and never used for single-variable claims.&lt;/p&gt;

&lt;p&gt;All three judges get an identical plain-JSON rubric, no vendor-specific structured output, so their scores stay comparable. Each grades two things per answer: factual correctness 0-100 against the reference, and a grounded true/false verdict: is every claim in the answer actually supported by the passages retrieval returned?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: the groundedness judge caught me before it caught any model. In its first version the judge received bare chunk texts, while the generator had seen the same passages labeled with rank and source filename. So judges kept marking legitimate "Sources:" citations as unsupported claims: 21 not-grounded verdicts, mostly artifact. After the&lt;br&gt;
fix (the judge now sees exactly what the generator saw) 10 remained, all real. If the judge's input differs from the generator's input, you are measuring the difference between the two prompts, not the model.&lt;/p&gt;

&lt;p&gt;Getting three new vendors into the run had its own friction: &lt;code&gt;Voyage&lt;/code&gt; rate-limits hard until a payment method exists, and &lt;code&gt;Gemini&lt;/code&gt;'s OpenAI-compatible embeddings endpoint returns no usage data and no batch indices, which my provider layer had to learn the hard way. Also worth saying out loud, given the thesis of the comparison piece: running this experiment handed my corpus and questions to three vendors that never had them before.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqtyq3gg988gaagwcm0o0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqtyq3gg988gaagwcm0o0.png" alt="Bar chart showing mean correctness scores across six RAG pipelines evaluated by three different judges" width="799" height="508"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Mean correctness per pipeline (answerable questions only), as three judges see it, listed as Gemini judge / GPT judge / Claude judge, then the measured generation cost per 100 answers and mean latency:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;p1 &lt;code&gt;oai-small&lt;/code&gt; + &lt;code&gt;Sonnet&lt;/code&gt;: 89.3 / 88.2 / 90.1; $0.58, 2.6s&lt;/li&gt;
&lt;li&gt;p2 &lt;code&gt;oai-small&lt;/code&gt; + &lt;code&gt;Haiku&lt;/code&gt;: 88.5 / 87.1 / 93.4; $0.22, 2.4s&lt;/li&gt;
&lt;li&gt;p3 &lt;code&gt;oai-small&lt;/code&gt; + &lt;code&gt;GPT Terra&lt;/code&gt;: 88.1 / 86.5 / 90.2; $0.33, 1.6s&lt;/li&gt;
&lt;li&gt;p4 &lt;code&gt;Voyage&lt;/code&gt; + &lt;code&gt;Sonnet&lt;/code&gt;: 93.1 / 90.5 / 90.9; $0.59, 2.8s&lt;/li&gt;
&lt;li&gt;p5 all-&lt;code&gt;Gemini&lt;/code&gt;: 94.3 / 93.4 / 93.6; $0.13, 2.0s&lt;/li&gt;
&lt;li&gt;p6 &lt;code&gt;oai-small&lt;/code&gt; + &lt;code&gt;DeepSeek&lt;/code&gt;: 87.4 / 85.7 / 88.6; $0.066, 1.1s&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What the matrix actually says:&lt;/p&gt;

&lt;p&gt;The corpus decides more than the model does. Of 80 answerable questions, 66 pass on all six pipelines and 3 fail on all six; only 11 depend on model choice at all. The largest spread between pipelines within one topic is 14 points, while the largest spread between topics within one pipeline is 20. Every pipeline finds licenses hard (topic means 74.6 to 81.5) and nutrition easy (93.0 to 99.5), for the same reasons.&lt;/p&gt;

&lt;p&gt;The embedder moved scores more than the generator. p4 against p1 is a pure embedder swap with the generator held fixed, and it is worth between 0.8 and 3.8 points depending on the judge; under every judge that beats what the Sonnet-to-Terra generator swap moves. The component that almost never appears in benchmarks turned out to be the biggest single-variable lever in this matrix.&lt;/p&gt;

&lt;p&gt;The two shared retrieval failures split cleanly. One question (the WHO saturated-fat limit, which exists verbatim in the corpus) scored 0 on every pipeline using the &lt;code&gt;OpenAI&lt;/code&gt; or &lt;code&gt;Voyage&lt;/code&gt; embedder and 100 on the &lt;code&gt;Gemini&lt;/code&gt; embedder: a genuine embedder-quality difference. The other (the BSD-3-Clause question) went unanswered on all six pipelines, because&lt;br&gt;
license texts never name themselves; the name lives only in the filename, which no chunk-text embedder can see. (16 of its 18 judge cells score it 0; the &lt;code&gt;Claude&lt;/code&gt; judge twice credited the refusal as correct, a small judge-behavior story of its own.) No&lt;br&gt;
model purchase fixes a corpus property. The fix for that class is contextual retrieval, measured in&lt;br&gt;
&lt;a href="https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049"&gt;the previous article&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Every model refused every unanswerable question. All six generators, &lt;code&gt;DeepSeek&lt;/code&gt; included, refused all 20&lt;br&gt;
questions whose answer is absent from the corpus, bait included, under all three judges. The refusal behavior comes from the prompt boundary, not from the model choice. Meaning, the cheapest generator in this matrix follows "answer only from context" as reliably as the most expensive one.&lt;/p&gt;

&lt;p&gt;Right for the wrong reason exists, and groundedness sees it. 16 of 1440 answerable verdicts came back&lt;br&gt;
not-grounded; 10 of those had scored 70 or above. The cleanest case: asked about charge sharing in a standard SCT, two pipelines answered correctly and scored 100, while the passage they retrieved states that principle for SCT Inst, a different scheme. The model knew the right answer from somewhere else. A correctness-only eval reports 100 and hides that the retrieval failed. Notably, zero of these cases happened on the two better embedders.&lt;/p&gt;

&lt;p&gt;No consistent self-preference. Each judge does nudge its own family up a little, but all three agree on the ends of the ranking, the all-&lt;code&gt;Gemini&lt;/code&gt; stack first and &lt;code&gt;DeepSeek&lt;/code&gt; last, and the &lt;code&gt;GPT&lt;/code&gt; judge actually scores its own family second-lowest. The one real disagreement: the &lt;code&gt;Claude&lt;/code&gt; judge alone ranks &lt;code&gt;Haiku&lt;/code&gt; first among generators, above &lt;code&gt;Sonnet&lt;/code&gt; itself; the other two judges keep &lt;code&gt;Sonnet&lt;/code&gt; ahead. Disagreement like this is a finding, so no number in this article merges scores from different judges.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: the whole experiment cost \$1.92 in generation, $0.07 in indexing, and \$9.57 in judging, metered at the config's list prices and re-checked against vendor pages before publishing. Two prices moved under my feet in those ten days: &lt;code&gt;OpenAI&lt;/code&gt; started a promo on the judge model, and &lt;code&gt;DeepSeek&lt;/code&gt; restructured its whole pricing into peak and off-peak tiers the day before the run, so p6's exact dollars are the least certain of the list (at most of the new&lt;br&gt;
rates it gets cheaper still). At this volume the drift is cents. The durable point: evaluating the answers cost five times more than producing them, which is worth knowing before someone promises to LLM-judge everything in a production budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Swapping the generator changed almost nothing. Six stacks from four vendors fail on the same questions and pass on the same ones; 66 of 80 outcomes are identical everywhere. &lt;code&gt;Haiku&lt;/code&gt; costs 37% of &lt;code&gt;Sonnet&lt;/code&gt;, &lt;code&gt;DeepSeek&lt;/code&gt; 11%, and at those discounts the answers stay in the same band.&lt;/li&gt;
&lt;li&gt;Which makes budget-routing the generator a real option: once your own eval has proven the slot interchangeable, a router like &lt;code&gt;OpenRouter&lt;/code&gt; can serve whatever is cheap this month, and the golden set catches any regression. I would not run the experiment itself through a router, since you never know which host actually served the tokens, but production serving after the experiment is exactly where one fits. Two costs to name: one more company sees your questions, and routers carry chat models only, so the embedder stays with its vendor either way.&lt;/li&gt;
&lt;li&gt;The embedder is the slot that still matters: up to 3.8 points between embedders with the generator fixed, and one question only the &lt;code&gt;Gemini&lt;/code&gt; embedder solved. Pin it, version it, give it its own regression test.&lt;/li&gt;
&lt;li&gt;The failures no model fixed are corpus properties, like a license text that never names itself. Those get fixed in the index, not in the model catalog.&lt;/li&gt;
&lt;li&gt;All six generators refused all 20 trap questions, so the refusal discipline lives in the prompt and survives model swaps with everything else.&lt;/li&gt;
&lt;li&gt;This is also my answer to vendor benchmark charts. On my 100 questions the ranking depends on which judge grades it and which topic you slice, and the gaps between generators are about the size of the judges' disagreement over them. A leaderboard saying some model is two points better tells you nothing about your corpus. Running your own hundred questions costs a morning and, in my case, $11.56 end to end.&lt;/li&gt;
&lt;li&gt;Still a small ruler: 20 answerable questions per topic, one run per cell, so one question is worth 5 points of a topic mean. Honest rather than optimized; treat the deltas, not the decimals.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Prompt sensitivity: one shared answer prompt was a design constant, so nothing here says how much the numbers move when the prompt does.&lt;/li&gt;
&lt;li&gt;Contextual retrieval on the matrix corpora. Production already ships it (previous article); the matrix ran the plain pipeline on purpose, to stay comparable with the framework comparison. Re-running the matrix on contextual indexes is the natural phase 2, and the filename-blindness result predicts what it fixes. Prepending source names to chunk text is the cheap half-step in the same direction.&lt;/li&gt;
&lt;li&gt;Serving through a router (the &lt;code&gt;OpenRouter&lt;/code&gt; point above) is argued here, not measured.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;: &lt;code&gt;comparison/matrix/&lt;/code&gt; (config, runners, golden sets), ADR-0021 (the experiment), ADR-0023 (contextual retrieval)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-1-framework-serhiy-kucherenko-tbdoe/" rel="noopener noreferrer"&gt;Comparing RAGs, part 1&lt;/a&gt; and &lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-2-benchmark-serhiy-kucherenko-kflse/" rel="noopener noreferrer"&gt;part 2&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049"&gt;Improving the scores of a RAG&lt;/a&gt; (contextual retrieval, the fix for the failure class no embedder solved)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/contextual-retrieval" rel="noopener noreferrer"&gt;Anthropic: Contextual Retrieval&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>python</category>
    </item>
    <item>
      <title>Improving the scores of a RAG</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Sat, 12 Sep 2026 11:28:33 +0000</pubDate>
      <link>https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049</link>
      <guid>https://dev.to/skucherenko/improving-the-scores-of-a-rag-5049</guid>
      <description>&lt;p&gt;In the previous articles (&lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-1-framework-serhiy-kucherenko-tbdoe/" rel="noopener noreferrer"&gt;part 1&lt;/a&gt; and &lt;a href="https://www.linkedin.com/pulse/comparing-rags-part-2-benchmark-serhiy-kucherenko-kflse/" rel="noopener noreferrer"&gt;part 2&lt;/a&gt;) I compared my own &lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;RAG&lt;/a&gt; against five other systems. One of its numbers bothered me since then: recall@5 = 0.60. In plain terms, for 4 questions out of 10 the page with the answer was actually retrieved, but ranked below the top 5 that the LLM gets to see. So the answers tended to avoid specifics ("settles instantly, see section 4.2.3" instead of "5 seconds"), and one question failed so badly that the judge scored the answer 0.&lt;/p&gt;

&lt;p&gt;The cause was measured earlier, and it is quite illustrative. The same page (p26 of the SCT Inst rulebook, the one that says "5 seconds") ranks differently depending on nothing but the phrasing of the question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"target maximum execution time", the spec's own words: rank 1&lt;/li&gt;
&lt;li&gt;"how fast does an SCT Inst payment settle?", a user's words: rank 9&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A user asks casually while the rulebook is written formally, so their embeddings end up far apart. The chunk size makes it worse: a ~300-word chunk embeds as an average of everything in it, meaning a single bare fact competes with the rest of its own chunk.&lt;/p&gt;

&lt;p&gt;There are two known ways to close this gap, so I measured both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;p&gt;A vocabulary mismatch can be attacked in two places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;At query time&lt;/strong&gt;: rewrite the casual question into several spec-vocabulary variants and search with all of them. Cheap, no re-indexing needed, and very popular (RAG tutorials call it multi-query or RAG-Fusion).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At index time&lt;/strong&gt;: give every chunk back the context it lost when it was cut out of its document, so that a bare "5 seconds" embeds as what it actually is. This is &lt;a href="https://www.anthropic.com/news/contextual-retrieval" rel="noopener noreferrer"&gt;Contextual Retrieval&lt;/a&gt; as described by Anthropic; they report large reductions in failed retrievals. Costs a one-time re-index.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: there is a third popular trick, &lt;a href="https://arxiv.org/abs/2212.10496" rel="noopener noreferrer"&gt;HyDE&lt;/a&gt;, where you embed a hypothetical&lt;br&gt;
&lt;em&gt;answer&lt;/em&gt; instead of the question. I rejected it without testing: in a compliance-heavy domain, a hallucinated draft steering the retrieval is exactly the failure mode this project exists to avoid.&lt;/p&gt;

&lt;p&gt;Both experiments follow the same rule as before: measure on the golden set first, and the production path changes only if the number justifies it. For reference, reranking passed through the same gate earlier: it lifted recall from 0.60 to 0.70 but costs about a minute per question, so it never shipped.&lt;/p&gt;
&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Fix 1: multi-query
&lt;/h3&gt;


&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q[casual question] --&amp;gt; RW["LLM rewriter&amp;lt;br/&amp;gt;(Haiku, ~$0.0001)"]
    RW --&amp;gt; V1[variant 1]
    RW --&amp;gt; V2[variant 2]
    RW --&amp;gt; V3[variant 3]
    Q --&amp;gt; R0[retrieve top-10]
    V1 --&amp;gt; R1[retrieve top-10]
    V2 --&amp;gt; R2[retrieve top-10]
    V3 --&amp;gt; R3[retrieve top-10]
    R0 --&amp;gt; F["reciprocal rank fusion&amp;lt;br/&amp;gt;(already in the codebase)"]
    R1 --&amp;gt; F
    R2 --&amp;gt; F
    R3 --&amp;gt; F
    F --&amp;gt; K[top-5 to the LLM]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;An LLM rewrites the question into spec-vocabulary variants, each variant retrieves separately, and the ranked lists are merged with the same reciprocal-rank-fusion function that hybrid search already used. Nothing hallucinated can leak in this way: whatever the phrasing, retrieval only ever returns real corpus text.&lt;/p&gt;

&lt;p&gt;The first surprise came before any quality result. Two identical eval runs returned 0.50 and 0.70. The reason was that the rewriter ran at the default temperature, so every run searched with different variants. If a retrieval mode has an LLM inside and the temperature is not pinned, the recall number is not reproducible. With temperature 0 the result became stable: 0.60, twice.&lt;/p&gt;

&lt;p&gt;Which is... exactly the baseline. No improvement. The hits did move though: multi-query recovered the currency question (the one behind the judge-0 answer) and lost a different one, one for one. That trade was the most useful outcome of the experiment, because it showed the vocabulary gap is real and reachable, and also that rewriting alone cannot cash it in. A better phrasing still lands on the same diluted chunk embeddings.&lt;/p&gt;

&lt;p&gt;I stopped there. With ten golden questions every hit is worth ±0.10, and tuning variant counts until the number goes up would be overfitting the eval, not improving retrieval.&lt;/p&gt;
&lt;h3&gt;
  
  
  Fix 2: contextual retrieval
&lt;/h3&gt;

&lt;p&gt;For contrast, this is what indexing looked like until now. Each chunk is embedded exactly as stored:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    PDF[rulebook PDF] --&amp;gt; CH[chunk ~300 words]
CH --&amp;gt; E["embed(chunk)"]
CH --&amp;gt; ST[store verbatim chunk]
E --&amp;gt; DB[(pgvector)]
ST --&amp;gt; DB&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;And with contextual retrieval, two LLM steps get inserted before the embedding. The same PDF feeds both paths: its full text produces a one-time summary, and that summary is the shared context for writing a short blurb per chunk:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    PDF[rulebook PDF] --&amp;gt; CH[chunk ~300 words]
PDF --&amp;gt;|full text| S["LLM summary&amp;lt;br/&amp;gt;(once per doc)"]
S --&amp;gt; B["LLM blurb per chunk&amp;lt;br/&amp;gt;(one call per page)"]
CH --&amp;gt; B
B --&amp;gt; E["embed(blurb + chunk)"]
CH --&amp;gt; ST[store verbatim chunk]
E --&amp;gt; DB[(pgvector)]
ST --&amp;gt; DB&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The blurb is one or two sentences situating each chunk: which rulebook it is from, which rule it belongs to, what its numbers are about. It is prepended only for the embedding. The stored text, the one that gets cited and shown as evidence, stays the verbatim spec passage. A bare "5 seconds" now embeds as "SCT Inst rulebook, target maximum execution time... 5 seconds", which is what a casual question is actually reaching for.&lt;/p&gt;

&lt;p&gt;Two implementation choices kept it cheap and safe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All chunks of a page get their blurbs in one call, against a one-time per-document summary. That is about $0.5 for the whole 484-chunk corpus, instead of the naive chunk-times-whole-document approach.&lt;/li&gt;
&lt;li&gt;If a blurb call fails, that page just embeds bare chunks with a warning instead of aborting the run. In practice it never happened: 484 of 484 chunks got contextualized.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flrnzp4edowvto461qkgl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flrnzp4edowvto461qkgl.png" alt=" " width="799" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Recall@5 per retrieval configuration, with the per-query cost of each:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;plain index (baseline): 0.60&lt;/li&gt;
&lt;li&gt;multi-query rewrites: 0.60, at ~1s + $0.0001 per query (and 0.50 to 0.70 until the temperature was pinned)&lt;/li&gt;
&lt;li&gt;reranking, from an earlier experiment, for scale: 0.70, at about a minute per query&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;contextual index: 0.80, at zero per-query cost&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;contextual index plus multi-query: 0.80, the hits only reshuffle, no net gain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The contextual lift is strictly additive: every question the baseline got right, plus two it did not, including the one behind the only judge-0 answer. It also carries through end to end. The answer eval on the new index scores 96.3 mean with a 100% pass rate, up from 84.8 and 90%. Asked live, the formerly failing question now answers "SCT Inst payments are made in euro" and cites the rulebook pages.&lt;/p&gt;

&lt;p&gt;One more detail: promoting the fix to production cost zero additional API dollars. The deploy tooling copies the chunks table, vectors included, from the local database to the cloud one, so production serves the exact index the numbers were measured on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: the earlier &lt;a href="https://github.com/KucherenkoSerhiy/payments-rag/blob/main/docs/comparison-report.md" rel="noopener noreferrer"&gt;six-system comparison&lt;/a&gt; scored this same system 84.8, fourth place out of six. On the new index it evals at 96.3, a hair under that comparison's winner (&lt;code&gt;LlamaIndex&lt;/code&gt;, 96.5). This is not a re-ranking of the comparison: different index, different day, and the other five systems would deserve the same fix applied. But it does suggest the gap was never about the framework choice. It was about the index.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The cheap query-side trick did nothing, and the index-side fix did everything. Multi-query is what every tutorial reaches for first because it needs no re-indexing; on this corpus it moved recall not at all. Contextual retrieval, for about $0.5 one time, moved it from 0.60 to 0.80 with zero per-query cost.&lt;/li&gt;
&lt;li&gt;The failed experiment still paid for itself. Its one gained question predicted what the real fix would deliver, and it surfaced that an unpinned LLM temperature makes an eval non-reproducible. A measured "no" is material, not waste.&lt;/li&gt;
&lt;li&gt;Ten questions is a small ruler. Every hit is worth ±0.10, so I stopped tuning the moment the temptation appeared. The numbers above are honest rather than optimized.&lt;/li&gt;
&lt;li&gt;Same discipline as always: nothing ships on vibes. Reranking measured well and stayed benched because of latency, multi-query measured flat and stayed benched, contextual retrieval measured +0.20 at zero query cost and shipped the same day.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Sentence-window / small-to-big chunking, the other index-side fix from the same playbook; it stacks with this one but was not measured here&lt;/li&gt;
&lt;li&gt;Re-running the full six-system comparison on contextual indexes (every system would need the same treatment to keep it fair)&lt;/li&gt;
&lt;li&gt;BM25/hybrid contextualization: the blurbs currently improve only the vector side&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;, ADR-0022 (multi-query), ADR-0023 (contextual retrieval), the retrieval-quality playbook&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/contextual-retrieval" rel="noopener noreferrer"&gt;Anthropic: Contextual Retrieval&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2212.10496" rel="noopener noreferrer"&gt;HyDE paper&lt;/a&gt;: Gao et al., "Precise Zero-Shot Dense Retrieval without Relevance Labels" (the hypothetical-answer trick, rejected here)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/KucherenkoSerhiy/payments-rag/blob/main/docs/comparison-report.md" rel="noopener noreferrer"&gt;The six-system comparison report&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>Comparing RAGs, Part 2: the benchmark</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Mon, 07 Sep 2026 20:23:13 +0000</pubDate>
      <link>https://dev.to/skucherenko/comparing-rags-part-2-the-benchmark-d06</link>
      <guid>https://dev.to/skucherenko/comparing-rags-part-2-the-benchmark-d06</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/skucherenko/comparing-rags-part-1-the-framework-e9p"&gt;Part 1&lt;/a&gt;, I laid out why "better" RAG isn't about raw accuracy: it's about whether your&lt;br&gt;
data stays inside your own infrastructure or ends up as a standing index on someone else's. Six approaches got tested&lt;br&gt;
head-to-head against the same 10-question golden set, &lt;code&gt;payments-rag&lt;/code&gt;'s own hand-rolled build, &lt;code&gt;openai-file-search&lt;/code&gt;,&lt;br&gt;
Google &lt;code&gt;NotebookLM&lt;/code&gt;, and three library builds (&lt;code&gt;Haystack&lt;/code&gt;, &lt;code&gt;LlamaIndex&lt;/code&gt;, &lt;code&gt;LangChain&lt;/code&gt;/&lt;code&gt;LangGraph&lt;/code&gt;). Here's how each was&lt;br&gt;
actually built, and what happened when they ran.&lt;/p&gt;

&lt;h2&gt;
  
  
  Development
&lt;/h2&gt;

&lt;p&gt;The source code lies in a &lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;.&lt;br&gt;
But the inner works across all six implementations can be summarized in the following sections.&lt;/p&gt;

&lt;h3&gt;
  
  
  System Design
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18yts03ii58400znl44l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18yts03ii58400znl44l.png" alt="System design: one test client drives all six systems into a shared adapter, judge and score" width="799" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This project consists of the following components:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Corpus&lt;/strong&gt;: the raw PDFs (three SEPA/ISO 20022 rulebooks), the input every system indexes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Indexing&lt;/strong&gt; (one-time, per system): chunk the corpus, embed each chunk, store as vectors.
&lt;code&gt;payments-rag&lt;/code&gt;/&lt;code&gt;Haystack&lt;/code&gt;/&lt;code&gt;LlamaIndex&lt;/code&gt;/&lt;code&gt;LangChain&lt;/code&gt; all do this themselves; &lt;code&gt;openai-file-search&lt;/code&gt; and &lt;code&gt;NotebookLM&lt;/code&gt; do
it
inside the vendor's own managed service instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embedding&lt;/strong&gt;: turns text (a question, or a chunk) into a vector. The same OpenAI model across every system except
&lt;code&gt;NotebookLM&lt;/code&gt;, whose embedding step is entirely internal/opaque.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval&lt;/strong&gt;: given a question's embedding, find the closest stored chunks. &lt;code&gt;pgvector&lt;/code&gt; similarity search for
&lt;code&gt;payments-rag&lt;/code&gt;; each framework's own equivalent for the rest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation&lt;/strong&gt;: question + retrieved chunks → an LLM → the answer. &lt;code&gt;Claude&lt;/code&gt; for &lt;code&gt;payments-rag&lt;/code&gt;, &lt;code&gt;gpt-4o&lt;/code&gt; for most
framework builds and &lt;code&gt;openai-file-search&lt;/code&gt;, &lt;code&gt;Gemini&lt;/code&gt; for &lt;code&gt;NotebookLM&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adapter&lt;/strong&gt; (comparison-only, not part of any system's real architecture): a thin wrapper per system normalizing
each one's very different calling convention into one common shape: question in, &lt;code&gt;{answer, contexts, cost, latency}&lt;/code&gt;
out. This is what makes the next layer possible across six otherwise-incompatible systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Judge/Eval&lt;/strong&gt; (offline, never per-query): the normalized output gets scored two ways: &lt;code&gt;RAGAS&lt;/code&gt; (
faithfulness/relevancy/precision/recall) and a cross-model judge (correctness vs. ground truth). Runs after the fact,
on the fixed 10-question golden set.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client vs. pipeline test client&lt;/strong&gt;: a real user's question only ever touches &lt;code&gt;payments-rag&lt;/code&gt;'s own &lt;code&gt;Embedder&lt;/code&gt; →
&lt;code&gt;Retriever&lt;/code&gt; → &lt;code&gt;Generator&lt;/code&gt; chain and never reaches eval. The pipeline test client (the golden-set harness) is the only
thing that reaches &lt;code&gt;Adapter&lt;/code&gt;. It's also the only thing that talks to the other five systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reaching the other five systems&lt;/strong&gt;: four of them (&lt;code&gt;openai-file-search&lt;/code&gt;, &lt;code&gt;Haystack&lt;/code&gt;, &lt;code&gt;LlamaIndex&lt;/code&gt;, &lt;code&gt;LangChain&lt;/code&gt;) have
a real API, so the test client calls them directly, same as it calls into &lt;code&gt;payments-rag&lt;/code&gt;. &lt;code&gt;NotebookLM&lt;/code&gt; doesn't, and
it's reached through a &lt;code&gt;Browser&lt;/code&gt; instead.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Indexing and Retrieving Text
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. &lt;code&gt;payments-rag&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7061p7jzqidgdqaqalbg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7061p7jzqidgdqaqalbg.png" alt="payments-rag sequence: embed, pgvector search, Claude answer" width="800" height="389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. &lt;code&gt;openai-file-search&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxn7jogzcibn6j1apqljk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxn7jogzcibn6j1apqljk.png" alt="openai-file-search sequence: one-time upload to a persistent OpenAI vector store, then per-question retrieval inside the API" width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Google &lt;code&gt;NotebookLM&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3nofb6u0db6djjs5xdm1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3nofb6u0db6djjs5xdm1.png" alt="NotebookLM sequence: manual upload and manual copy-paste per question" width="800" height="479"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. &lt;code&gt;Haystack&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz3uh80fl1nesblhpowha.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz3uh80fl1nesblhpowha.png" alt="Haystack sequence: PyPDF load, split, embed into in-memory store, retrieve top 5, generate" width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxspjdlr00rtn7fzd5o1i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxspjdlr00rtn7fzd5o1i.png" alt="Haystack bug: one malformed 12,000-word 'sentence' exceeds the embedding size limit and the chunk is silently dropped" width="800" height="133"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Haystack&lt;/code&gt;'s sentence-based chunker treated a malformed PDF-extracted block as one giant "sentence," pushing it past&lt;br&gt;
OpenAI's embedding size limit. As a consequence, &lt;code&gt;Haystack&lt;/code&gt; silently dropped that chunk instead of raising an error,&lt;br&gt;
quietly losing a third of the index. Fixed by switching to &lt;code&gt;Haystack&lt;/code&gt;'s own word-based chunking default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. &lt;code&gt;LlamaIndex&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgnxaqgva1vdmucd9dyi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgnxaqgva1vdmucd9dyi.png" alt="LlamaIndex sequence: pypdf load, sentence split, VectorStoreIndex, query engine" width="800" height="651"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc0dwaqt78i0lj3px3o8l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc0dwaqt78i0lj3px3o8l.png" alt="LlamaIndex bug: raw PDF metadata and binary noise leaked into extracted text, outranking real answers" width="799" height="148"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;LlamaIndex&lt;/code&gt;'s default PDF reader leaked raw PDF metadata into extracted text and mangled some pages into unreadable&lt;br&gt;
binary noise, both silently, without an error thrown. The noise sometimes outranked the real answer during retrieval.&lt;br&gt;
Fixed by loading PDFs with &lt;code&gt;pypdf&lt;/code&gt; directly, the same library &lt;code&gt;Haystack&lt;/code&gt; already used cleanly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. &lt;code&gt;LangChain&lt;/code&gt; / &lt;code&gt;LangGraph&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F802owkpjhuvevwht7e82.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F802owkpjhuvevwht7e82.png" alt="LangChain/LangGraph sequence: pypdf load, recursive splitter, retrieve and generate nodes in a compiled graph" width="800" height="494"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznjm7mkyhr8kmcye4aqn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznjm7mkyhr8kmcye4aqn.png" alt="Results: judge score, RAGAS metrics, cost and latency for all six systems" width="800" height="681"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;LlamaIndex&lt;/code&gt; is the best in both judge and faithfulness scores.&lt;/p&gt;

&lt;p&gt;OpenAI file search cost $0.3526 for the same ten questions, which is the most expensive system in the comparison&lt;br&gt;
by a wide margin: over 8x pricier than the next-priciest option (&lt;code&gt;Haystack&lt;/code&gt;, $0.0413), and roughly 13x pricier than&lt;br&gt;
&lt;code&gt;payments-rag&lt;/code&gt;'s own hand-rolled approach ($0.0276). Likely because it stuffs more retrieved context into every request&lt;br&gt;
than expected.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;NotebookLM&lt;/code&gt; costs nothing in dollars, but it required 11 min of setup work of clicking&lt;br&gt;
before the first question could even be asked. Meaning, it just moves the cost from money to time.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;NotebookLM&lt;/code&gt; ties &lt;code&gt;LlamaIndex&lt;/code&gt; for the highest answer-relevancy of all six, despite its other &lt;code&gt;RAGAS&lt;/code&gt; metrics being&lt;br&gt;
unusable. &lt;code&gt;LangChain&lt;/code&gt;, the most popular framework by mindshare, placed second-worst, just above &lt;code&gt;Haystack&lt;/code&gt;.&lt;br&gt;
&lt;code&gt;Haystack&lt;/code&gt;'s remaining weakness is a retrieval bias toward one source document on standard-SCT questions, not a&lt;br&gt;
grounding failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Some numbers worth pulling forward before the caveats:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;own built &lt;code&gt;payments-rag&lt;/code&gt; placed 4th of 6 on judge-scored accuracy (84.8), behind &lt;code&gt;LlamaIndex&lt;/code&gt; (96.5),
&lt;code&gt;NotebookLM&lt;/code&gt; (92.5), and &lt;code&gt;openai-file-search&lt;/code&gt; (90.2).&lt;/li&gt;
&lt;li&gt;it was also the cheapest ($0.0276) and fastest (2.68s) system in the whole comparison. For instance, OpenAI markup
caused it to cost about 8-13x, while &lt;code&gt;NotebookLM&lt;/code&gt; took 11 minutes to set up.&lt;/li&gt;
&lt;li&gt;silent bugs may creep simply because of how the systems are wired up as it happened to &lt;code&gt;LlamaIndex&lt;/code&gt; and &lt;code&gt;Haystack&lt;/code&gt;. In
this case it was about clean corpus and a sane chunk size.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The obvious flaw of this stage is a small set of questions (or maybe topics of questions) of a golden set.&lt;br&gt;
Additionally, the compared options could be expanded dimension-wise by combining different models&lt;br&gt;
(both embedding, llm, judge).&lt;/p&gt;

&lt;p&gt;However, this intel brings some clarity on why we might end up having RAGs everywhere as well as concerns about both&lt;br&gt;
correctness and security. Building a working and useful RAG is one effort. Building one at scale is a different thing.&lt;br&gt;
And building one for privacy is yet another problem. The only way to fully avoid exposure is self-hosting all&lt;br&gt;
the layers: the embedding, the generation models, and the retrieval layer. &lt;code&gt;payments-rag&lt;/code&gt;'s own hand-rolled setup still&lt;br&gt;
sends chunks to OpenAI and Anthropic at embed- and generation-time (see the exposure diagram in &lt;a href="https://dev.to/skucherenko/comparing-rags-part-1-the-framework-e9p"&gt;Part 1&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Out of Scope
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Bigger and more diverse golden sets, additional metrics beyond &lt;code&gt;RAGAS&lt;/code&gt;/judge&lt;/li&gt;
&lt;li&gt;Testing every embedding/LLM/judge model combination for full experimental cleanliness: this run fixed one model set
throughout&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;NotebookLM&lt;/code&gt;'s paid Enterprise tier: investigated, doesn't change the conclusion (still no way to ask it a question
programmatically)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also, that was made on Python. I'd imagine someone doing it, say, on C++ or Rust (whatever is better) would make it&lt;br&gt;
faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/vibrantlabsai/ragas/issues/2753" rel="noopener noreferrer"&gt;RAGAS issue #2753&lt;/a&gt;: the scoring-library bug&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.cloud.google.com/gemini/enterprise/notebooklm-enterprise/docs/api-notebooks" rel="noopener noreferrer"&gt;NotebookLM Enterprise API docs&lt;/a&gt;:
no query endpoint, any tier&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://support.google.com/notebooklm/answer/17004255" rel="noopener noreferrer"&gt;NotebookLM privacy policy&lt;/a&gt;: training-data use&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.haystack.deepset.ai/docs/telemetry" rel="noopener noreferrer"&gt;Haystack telemetry docs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>Comparing RAGs, Part 1: the framework</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Sat, 05 Sep 2026 20:25:59 +0000</pubDate>
      <link>https://dev.to/skucherenko/comparing-rags-part-1-the-framework-e9p</link>
      <guid>https://dev.to/skucherenko/comparing-rags-part-1-the-framework-e9p</guid>
      <description>&lt;p&gt;In the past, I built my own &lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;RAG&lt;/a&gt;. The idea was to implement and debug concepts rather than read documentation. So far, a RAG is supposed to address this issue for a person or a company:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AS an engineer,
I WANT TO get accurate data on need
SO THAT I get the exact answer, with a citation, as fast as possible
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the problem a RAG tries to solve. Since this is an engine that also includes a database, safety and scalability concerns apply. There are multiple options, and the problem might be: &lt;strong&gt;which option is the best?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Hypothesis
&lt;/h2&gt;

&lt;p&gt;We could categorize solutions by usage and integration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Manual personal interaction services Google's like &lt;code&gt;NotebookLM&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;API services like OpenAI's &lt;code&gt;file_search&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Libraries like &lt;code&gt;Haystack&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Custom solution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Typically, if a system needs scale (imagine lots of people constantly surfing through documentation which could be indexed in a RAG), the choice lies between building an own solution using a library and a cloud solution from another company.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning&lt;/strong&gt;: a claim like "no need for expertise because we have AI" means we have no way to verify if a clue detail was&lt;br&gt;
missed (described at &lt;a href="https://www.linkedin.com/pulse/judge-strictly-avoid-hallucinations-serhiy-kucherenko-cjsye/" rel="noopener noreferrer"&gt;Judge Strictly Avoid Hallucinations&lt;/a&gt;).&lt;/p&gt;
&lt;h3&gt;
  
  
  Safety and Politics as concerns
&lt;/h3&gt;

&lt;p&gt;Recent worldwide events have shown that depending on another company is not only a security vulnerability but also a political one. This matters because while it might be easier and work better initially to just dump everything into a cloud solution of another enterprise, this creates a single point of failure in many ways. Also tends to be quite expensive in the long run.&lt;/p&gt;

&lt;p&gt;Maintaining an own solution is not cheap either: you will have an internal service with a team maintaining it, and with the pace the AI tech progresses, you risk falling behind already in a few months if not weeks.&lt;/p&gt;
&lt;h3&gt;
  
  
  Define "better"
&lt;/h3&gt;

&lt;p&gt;So... &lt;strong&gt;which solution would work better, and how a RAG would benefit us?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before comparing anything, it's worth being precise about what "better" even means here. It isn't rather who answers questions most accurately, as we will end up just comparing frontier models and context windows on a ten-question set. The real question is &lt;strong&gt;integration&lt;/strong&gt;: can this run inside a system where data is &lt;strong&gt;private&lt;/strong&gt;?&lt;/p&gt;
&lt;h4&gt;
  
  
  The Judge
&lt;/h4&gt;

&lt;p&gt;The most helpful way is to use another &lt;a href="https://www.evidentlyai.com/llm-guide/llm-as-a-judge" rel="noopener noreferrer"&gt;LLM as a judge&lt;/a&gt; with strict scoring, to avoid any false positives, as these are the most dangerous silent killers of confidence. So a &lt;code&gt;gpt-4o&lt;/code&gt; with &lt;code&gt;text-embedding-3-small&lt;/code&gt; will do just fine as a different vendor from the &lt;code&gt;Claude&lt;/code&gt;-generated answers this project produces, which avoids grading its own homework.&lt;/p&gt;
&lt;h4&gt;
  
  
  RAGAS
&lt;/h4&gt;

&lt;p&gt;We will also use [&lt;code&gt;RAGAS&lt;/code&gt;] &lt;a href="https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/" rel="noopener noreferrer"&gt;https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/&lt;/a&gt;) 0.4.3 for evaluation. &lt;code&gt;RAGAS&lt;/code&gt; was picked as the scoring method for the comparison itself specifically because it is a library, in-process, introduces no new recipient of corpus data, and has a documented telemetry kill switch.&lt;/p&gt;

&lt;p&gt;This standard quartet of metrics will do:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Faithfulness&lt;/code&gt;: does the answer only say things the retrieved chunks actually support?&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ResponseRelevancy&lt;/code&gt;: does the answer address the question asked, not something adjacent?&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LLMContextPrecisionWithoutReference&lt;/code&gt;: of what got retrieved, how much was actually relevant?&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LLMContextRecall&lt;/code&gt;: of what should have been retrieved, how much actually was?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: even the scoring library itself shipped broken. A plain &lt;code&gt;pip install ragas&lt;/code&gt; crashes on import with &lt;code&gt;ModuleNotFoundError: No module named 'langchain_community.chat_models.vertexai'&lt;/code&gt;, caused by &lt;code&gt;RAGAS&lt;/code&gt;'s own code imports a path that newer versions of one of its own dependencies removed. Real, still-open bug on &lt;code&gt;RAGAS&lt;/code&gt;'s side&lt;br&gt;
(&lt;a href="https://github.com/vibrantlabsai/ragas/issues/2753" rel="noopener noreferrer"&gt;issue #2753&lt;/a&gt;, fix still unmerged), not something this project did wrong. Worked around by pinning an older &lt;code&gt;langchain-community&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;
  
  
  Issue: Transient vs. Persistent Exposure
&lt;/h3&gt;

&lt;p&gt;That's a second axis worth noting. &lt;em&gt;Transient&lt;/em&gt; means a request goes out, a vendor's model processes it, and nothing about your corpus survives on their infrastructure afterward as a queryable asset.&lt;/p&gt;

&lt;p&gt;Persistent means the opposite: your corpus (or an embedding of it) gets stored on someone else's infrastructure as a standing index that outlives the call that created it and keeps being a breach target indefinitely. OpenAI's &lt;code&gt;file_search&lt;/code&gt; and Google's &lt;code&gt;NotebookLM&lt;/code&gt; both fall into the second category: uploading your documents creates a managed index that exists independently of any single question you ask. That exposure means your private data may be breached or even secretly used as training by host companies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: a RAG already accepts some transient exposure, as we will end up sending a chunk (100-1000 words, which is a typical size of a chunk).&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;sequenceDiagram
    participant You as Your infrastructure
    participant Embed as Embedding model (cloud API)
    participant LLM as Answering LLM (cloud API)
    Note over You, LLM: Embed-time (chunk text leaves your infra)
    You -&amp;gt;&amp;gt; Embed: chunk text (index + query)
    Embed --&amp;gt;&amp;gt; You: embedding vector
    Note over You: stored/searched locally (Postgres/pgvector)
    Note over You, LLM: Generation-time (chunk text leaves your infra again)
    You -&amp;gt;&amp;gt; LLM: chunk text (as prompt context)
    LLM --&amp;gt;&amp;gt; You: answer&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The only way to avoid exposing your data completely is to host the embedding, the answering model, and the judging model yourself, at a cost of having weaker models and continuous maintenance.&lt;/p&gt;

&lt;h3&gt;
  
  
  Issue: Integration
&lt;/h3&gt;

&lt;p&gt;Another issue is how easy it is to integrate a RAG into your knowledge base. Typically, a team has a separate ticketing platform, data, and a logging storage. Also, it has a company-wide context, dependencies on other teams, documentation, alerts. That presents a set of challenges when trying to find a necessary set of pieces of info. As well as an issue with integrating such a solution as RAG and maintaining the relevance of its data.&lt;/p&gt;

&lt;p&gt;Besides, each team will have their unique and individual stuff like specified above. Meaning, having a single same corpus for all teams is not possible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: As of August 2026 &lt;code&gt;NotebookLM&lt;/code&gt; only works by hand, through its own web page, there's no way to script it or plug it into anything else&lt;br&gt;
even on the paid enterprise tier. Google added an API in 2025, but it still can't be asked a question through it. Regardless, given the pace AI tech evolves, it is still included in this comparison to keep an eye on.&lt;/p&gt;

&lt;h3&gt;
  
  
  The six systems compared, and how each was queried
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;payments-rag&lt;/code&gt;&lt;/strong&gt;: the project's own production path, &lt;code&gt;Claude&lt;/code&gt;-generated answers over its own &lt;code&gt;pgvector&lt;/code&gt; retriever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;openai-file-search&lt;/code&gt;&lt;/strong&gt;: OpenAI Responses API, &lt;code&gt;model="gpt-4o"&lt;/code&gt;,
&lt;code&gt;tools=[{"type": "file_search", "vector_store_ids": [...]}]&lt;/code&gt;,
corpus PDFs uploaded once to a named OpenAI vector store&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;NotebookLM&lt;/code&gt;: queried through its web UI, answers and citation filenames captured by hand&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Haystack&lt;/code&gt;&lt;/strong&gt;: a RAG pipeline built with the deepset library instead of hand-rolling it with the same corpus and
models. It is the framework that does the orchestration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: &lt;code&gt;Haystack&lt;/code&gt; ships with anonymous, opt-out telemetry (on by default), though it sends only component types (e.g., which retriever/store you used), no personal data. To turn it off: &lt;code&gt;HAYSTACK_TELEMETRY_ENABLED=False&lt;/code&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;LlamaIndex&lt;/code&gt;&lt;/strong&gt;: the same idea, built with a different library, included to see whether "framework vs. framework" matters as much as "framework vs. hand-rolled"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;LangChain&lt;/code&gt;&lt;/strong&gt;/&lt;strong&gt;&lt;code&gt;LangGraph&lt;/code&gt;&lt;/strong&gt;: a RAG pipeline built with &lt;code&gt;LangChain&lt;/code&gt;'s components (loader, splitter, vector store, retriever),
with &lt;code&gt;LangGraph&lt;/code&gt; doing the orchestration instead of a hand-rolled function or the other frameworks' own pipeline objects&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The comparison was based on answers of the same &lt;strong&gt;10-question golden set&lt;/strong&gt; that was used throughout the development of &lt;code&gt;payments-rag&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Part 2 covers how each of these was actually built, and what broke along the way.&lt;/p&gt;

&lt;p&gt;Publishing Monday.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to hurt yourself with AI</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Tue, 25 Aug 2026 14:56:39 +0000</pubDate>
      <link>https://dev.to/skucherenko/how-to-hurt-yourself-with-ai-4jp1</link>
      <guid>https://dev.to/skucherenko/how-to-hurt-yourself-with-ai-4jp1</guid>
      <description>&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;34x more work than it should have been.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timeline
&lt;/h2&gt;

&lt;p&gt;I write technical articles, and I like experimenting with something new in each one. For one of them, the experiment was an AI pipeline: it would draft the piece and pass it through several rounds of automated review before I saw it.&lt;/p&gt;

&lt;p&gt;That pipeline produced &lt;strong&gt;1,948&lt;/strong&gt; words while processed &lt;strong&gt;65,930&lt;/strong&gt; in total which is &lt;strong&gt;~34x&lt;/strong&gt; more.&lt;/p&gt;

&lt;p&gt;The idea was to base the writing on careful review and a feedback loop. So I ended up reviewing in person and wasting 34x words/tokens more than what I could actually produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Root cause
&lt;/h2&gt;

&lt;p&gt;I am not sure what the exact cause is. There are several candidate explanations, and I do not think any one of them is complete on its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Not enough context to judge whether a cut was a real improvement.&lt;/li&gt;
&lt;li&gt;The AI outputs what is locally coherent, what "sounds right," with no guarantee that coherent and applicable are the same thing.&lt;/li&gt;
&lt;li&gt;The AI was built to fill in gaps rather than flag them, and what it fills in with is often an assumption, usually one that turns out false.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What a human has that the AI does not is a felt sense of consequence: what a decision actually costs or breaks, in a specific world, learned from having been wrong before. "Experience" is the word for that, and it is overused to the point of meaning nothing, but I do not have a better one.&lt;/p&gt;

&lt;p&gt;If I compress all of that into one line: &lt;strong&gt;the pipeline optimized for the wrong thing.&lt;/strong&gt; It optimized for passing a check. It could not optimize for "is this piece actually good," because nothing in it could feel the cost of being wrong.&lt;/p&gt;

&lt;p&gt;The aim behind building it was sound and I would keep it: automate whatever a computer can do so a human does not have to. What failed was assuming editorial judgment was one of those things.&lt;/p&gt;

&lt;h2&gt;
  
  
  The debt
&lt;/h2&gt;

&lt;p&gt;Engineers already track several kinds of debt that are not financial.&lt;br&gt;
One of them is getting real attention now: &lt;strong&gt;cognitive debt&lt;/strong&gt;, which is what you take on when you outsource tracking why your own system does what it does to an LLM. Two recent sources define it the same way: a&lt;br&gt;
&lt;a href="https://arxiv.org/abs/2506.08872" rel="noopener noreferrer"&gt;MIT Media Lab study&lt;/a&gt; on AI-assisted writing, and &lt;a href="https://getdx.com/blog/cognitive-debt-the-hidden-risk-in-ai-driven-software-development/" rel="noopener noreferrer"&gt;getdx.com's framing&lt;/a&gt; for software teams specifically.&lt;/p&gt;

&lt;p&gt;This pipeline is a worked example. I could tell you the ratio, the gate verdicts, the timeline. For a while I could not have told you, from memory, what any single piece actually argued.&lt;/p&gt;

&lt;p&gt;The more AI gets pushed into everything, the more this particular kind of debt matters. On a prototype, nobody bothers tracking it. On something meant to hold up, not understanding what you built is a deal-breaker,&lt;br&gt;
and "just ship it, iterate later" is the same old fallacy wearing an&lt;br&gt;
AI-shaped coat: speed and understanding traded off as if they were opposites, when the trade was optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes
&lt;/h2&gt;

&lt;p&gt;The intent behind the pipeline was to be the reviewer: read a draft,&lt;br&gt;
judge it, hand back a correction, and only step in myself when the correction failed too. In practice, I was correcting constantly, sentence by sentence, round after round, and each correction only produced more issues.&lt;/p&gt;

&lt;p&gt;That does not work, and it should not. Scale it up and the failure gets clear: imagine ten thousand articles like this shipping out daily, all with a human correcting sentence by sentence, all costing that person a 34x reading tax.&lt;/p&gt;

&lt;p&gt;Going forward, AI in this workflow does small, bounded jobs: formatting, a lookup, a narrow research question with a checkable answer. Things that do not load much cognitively.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question for you
&lt;/h2&gt;

&lt;p&gt;What have you noticed, using AI or watching it get used badly? Two data points I found while researching this piece:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.livescience.com/technology/artificial-intelligence/i-violated-every-principle-i-was-given-ai-agent-deletes-companys-entire-database-in-9-seconds-then-confesses" rel="noopener noreferrer"&gt;A coding agent deleted a production database and its backups in nine seconds after it decided, on its own, to work around a credential mismatch it hit mid-task. Its own log afterward: "I violated every principle I was given."&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.damiencharlotin.com/hallucinations/" rel="noopener noreferrer"&gt;A public tracker has now logged over 1,200 cases of lawyers submitting AI-hallucinated fake case citations to real courts, adding new cases at roughly five to six a day.&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Different domain, same shape as this piece: a check that looked like it was working, until someone read closely enough to notice it wasn't.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Judge strictly to avoid hallucinations</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Fri, 14 Aug 2026 13:13:20 +0000</pubDate>
      <link>https://dev.to/skucherenko/judge-strictly-to-avoid-hallucinations-7kd</link>
      <guid>https://dev.to/skucherenko/judge-strictly-to-avoid-hallucinations-7kd</guid>
      <description>&lt;p&gt;Ten questions over the SEPA rulebooks. A second model grades each answer against a reference, zero to a hundred, and the run came back healthy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6k2g0q2y7drnzudw9ber.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6k2g0q2y7drnzudw9ber.png" alt="Ten questions graded 0 to 100. Nine cluster between 80 and 100. Question five sits flat at zero." width="800" height="284"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Except for question five, which asked which currency SCT Inst payments are executed in. The reference answer is one word long.&lt;/p&gt;

&lt;p&gt;A zero on a one-word question reads like a system that does not know the first thing about its own subject matter. That is not what happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  It was doing what it was told
&lt;/h2&gt;

&lt;p&gt;The production prompt is explicit about the case where retrieval comes up short:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Answer the question using ONLY the sources below. If they do not contain the&lt;br&gt;
answer, say so.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Retrieval never put the euro page in the top five results. So the sources handed to the model genuinely did not contain the answer, and the model was under written instruction to say exactly that.&lt;/p&gt;

&lt;p&gt;The grader took that response, compared it to the word "Euro", and returned a zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  The grader cannot see the difference
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg4qr3kcl9i0rukzxkszr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg4qr3kcl9i0rukzxkszr.png" alt="The judge receives question, reference answer and answer text, and returns one score. It never receives the retrieved sources, so a wrong guess and a correct refusal both land at zero." width="799" height="361"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The judge receives the question, the reference answer and the answer text. It does not receive the sources, and that is deliberate: a grader holding the retrieved chunks would end up grading the retrieval instead of the answer.&lt;/p&gt;

&lt;p&gt;That choice has a cost. "Correct, given what it was handed" is not a judgement this grader can make. It has one axis, factual match against the reference, and everything that fails to match lands at the bottom together.&lt;/p&gt;

&lt;p&gt;Which means a fabricated answer about euro payments would have scored exactly what the refusal scored. Zero, either way. The eval could not tell the two apart.&lt;/p&gt;

&lt;p&gt;The saved run does not help either. Each question is stored as an id, a score and a one-line critique. The answer text itself is not kept, so nothing downstream can recover the distinction the grader could not draw.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure was upstream
&lt;/h2&gt;

&lt;p&gt;If the refusal had been the problem, changing the answering model would have moved the number. It was not, and it did not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcr4uhva22q7p3rhk1zfp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcr4uhva22q7p3rhk1zfp.png" alt="The same question under three retrieval configurations. Dense vectors miss the euro page and the answer scores 0. Hybrid keyword search retrieves it but leaves recall@5 unchanged at 0.60, trading away a question the vectors had right. The eval-only reranker retrieves it and lifts recall@5 to 0.70." width="800" height="372"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two separate experiments recovered that question, and neither one touched generation. Keyword search found the euro page instantly, because the page contains the actual word. Vector search matches on meaning rather than&lt;br&gt;
words, and it ranked that page outside the top five. A cross-encoder reranker, kept to the eval path only, found it as well and lifted recall@5 from 0.60 to 0.70.&lt;/p&gt;

&lt;p&gt;Neither is a free fix. Leaving keyword search on permanently left recall@5 exactly where it started, at 0.60, because it dropped a question the vectors had been getting right. The blind spot moved. It did not close. The reranker is the only arm that moved the number, and it is not in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make a wrong answer cost something
&lt;/h2&gt;

&lt;p&gt;Retrieval will keep missing pages. Improving that is a separate job, so what matters here is how the eval reacts when it happens.&lt;/p&gt;

&lt;p&gt;The grader was not being harsh. My rubric just stops at zero, so both failures pile up there together. One of them wastes the reader's time. The other sends them into a payments rulebook holding a rule that does not exist. My eval scored those the same.&lt;/p&gt;

&lt;p&gt;So the strictness belongs on the wrong answer. &lt;a href="https://arxiv.org/abs/2406.04744" rel="noopener noreferrer"&gt;Meta's CRAG&lt;br&gt;
benchmark&lt;/a&gt;, which I only found after this run, grades four outcomes instead of one axis: a correct answer scores 1, a useful answer with minor errors 0.5, a &lt;strong&gt;missing&lt;/strong&gt; answer 0, and an incorrect answer &lt;strong&gt;minus one&lt;/strong&gt;. "Missing" is defined there as the system replying "I don't know".&lt;/p&gt;

&lt;p&gt;That minus sign is the part I was missing. Saying nothing costs you nothing. Being wrong costs you a point.&lt;/p&gt;

&lt;p&gt;Nothing was fabricated in this run. The system declined, as instructed, and the number punished it for that. But the rubric that could not credit that refusal is the same rubric that would not have charged for a made-up answer, and a made-up answer is the one that reaches the reader.&lt;/p&gt;

&lt;p&gt;Charge the model for being wrong. Let it say nothing for free.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The illusion of improvement</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Sun, 09 Aug 2026 19:24:47 +0000</pubDate>
      <link>https://dev.to/skucherenko/the-illusion-of-improvement-1j54</link>
      <guid>https://dev.to/skucherenko/the-illusion-of-improvement-1j54</guid>
      <description>&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://rag.serhiykucherenko.dev" rel="noopener noreferrer"&gt;My RAG&lt;/a&gt; answers questions about SEPA payment rulebooks. Every page of those PDFs carries the same header and footer, and all of it lands inside the chunks that get embedded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fluuhdq4hlyxxny15ouyc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fluuhdq4hlyxxny15ouyc.png" alt="One chunk, as the embedder saw it" width="800" height="194"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two lines of letterhead on every one of 290 pages, sitting in the same vector as the sentence that actually answers something.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hypothesis
&lt;/h2&gt;

&lt;p&gt;Identical text in every chunk means a shared component in every embedding. I believed that component was flattening the similarity signal: distances on my test query all clustered tightly around 0.34 instead of spreading out.&lt;/p&gt;

&lt;p&gt;Strip the repeated lines, and while I am in there, stop splitting mid-sentence. Sharper chunks, sharper vectors, better ranking.&lt;/p&gt;

&lt;p&gt;It is a mechanism, and it is checkable. I would have bet on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What improved
&lt;/h2&gt;

&lt;p&gt;The cleanup worked exactly as designed. Any line repeating on at least half a document's pages gets stripped, with digits masked first so page numbers cannot disguise a repeat. "&lt;a href="http://www.epc-cep.eu" rel="noopener noreferrer"&gt;www.epc-cep.eu&lt;/a&gt; 35" on one page and "&lt;a href="http://www.epc-cep.eu" rel="noopener noreferrer"&gt;www.epc-cep.eu&lt;/a&gt; 36" on the next both normalise to "&lt;a href="http://www.epc-cep.eu" rel="noopener noreferrer"&gt;www.epc-cep.eu&lt;/a&gt; #", counted once per page. That line recurs on all 290 pages, so it goes. The rulebook title and the "Date issued" line collapse the same way and go with it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Boilerplate in chunk text:&lt;/strong&gt; every chunk, to none.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Words on a typical page:&lt;/strong&gt; 464, to 448.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunks in the corpus:&lt;/strong&gt; 495, to 484.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunk boundaries:&lt;/strong&gt; cut mid-sentence, to whole sentences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Eleven chunks disappeared, and the reason is duller than it sounds. No chunk was pure boilerplate. The junk was about sixteen words per page, spread across every chunk on that page. Chunking runs per page at a 300-word target, so the only pages whose count changed were the eleven sitting just above that line: 315 words became 299, 306 became 290, and two chunks became one.&lt;/p&gt;

&lt;p&gt;Everywhere else the text simply got cleaner without changing shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reality check
&lt;/h2&gt;

&lt;p&gt;Then I re-ran the query I had been watching. &lt;strong&gt;Same top pages, same order.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The one thing I could have seen with my own eyes did not happen. Nothing reordered.&lt;/p&gt;

&lt;p&gt;That is the entire result, and it is worth being blunt about how weak an observation it is. There was no evaluation set yet. The test was me reading one query's results and forming an impression, and an impression cannot separate "no effect" from "an effect too small to notice".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the mechanism was wrong
&lt;/h2&gt;

&lt;p&gt;Ranking does not compare chunks to each other. It compares each chunk to the query and sorts by that angle. The boilerplate was only ever in the chunks; the question you type carries none of it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F337jc5aap914rowxyvgy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F337jc5aap914rowxyvgy.png" alt="The chunks moved. The order they came back in did not." width="800" height="431"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I have two explanations and they both fit. The junk was identical in every chunk, so removing it nudged them all in a similar direction instead of spreading them apart. It was also only about 3% of a page's words, which is not much of a shove to begin with. Either way, I predicted a reshuffle and got a nudge.&lt;/p&gt;

&lt;p&gt;What I cannot do is tell those two apart. Separating them needs the per-result distances to more than two decimal places, and I never wrote them down. Even a 3% shift should have moved something in the third decimal. Whether it did is a question my own logs cannot answer.&lt;/p&gt;

&lt;p&gt;The 0.34 floor was not noise the boilerplate added either. That is what a dense, homogeneous legal corpus looks like to an embedding model. The chunks really are that similar to each other, because rulebook pages really are that similar.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I kept, and what I changed
&lt;/h2&gt;

&lt;p&gt;The cleanup stayed. Clean chunks are what a citation displays against an exact page, and what the prompt feeds the model. The payoff is real. It is just not the payoff I built it for.&lt;/p&gt;

&lt;p&gt;What actually changed was my process. Three days later I built a golden set of questions with verified answers, and it produced a retrieval baseline of &lt;strong&gt;recall@5 = 0.60&lt;/strong&gt;, the number every retrieval change has had to argue with since.&lt;/p&gt;

&lt;p&gt;One detail I have left alone deliberately. The comment at the top of the cleanup module still states the dead hypothesis as fact: boilerplate "drags every embedding toward the same noise, which flattens the similarity signal". It is wrong, it is still there, and it is a better reminder than anything I would write on purpose.&lt;/p&gt;

&lt;p&gt;The hypothesis was reasonable, the mechanism was checkable, and the check said no. The next fix that sounds this obviously right gets measured before it gets believed.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Three small lessons from building a RAG by hand</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Sun, 02 Aug 2026 19:16:41 +0000</pubDate>
      <link>https://dev.to/skucherenko/three-small-lessons-from-building-a-rag-by-hand-43n0</link>
      <guid>https://dev.to/skucherenko/three-small-lessons-from-building-a-rag-by-hand-43n0</guid>
      <description>&lt;p&gt;A retrieval system over the SEPA payment rulebooks. Plain Python, Postgres, two LLM vendors. Three independent decisions from it are worth writing down, because in each case the decision was cheap and the thing it taught was not the thing it was made for.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The vector database never got installed
&lt;/h2&gt;

&lt;p&gt;Every RAG guide starts in the same place. &lt;a href="https://docs.langchain.com/oss/python/langchain/rag" rel="noopener noreferrer"&gt;LangChain's own documentation&lt;/a&gt;&lt;br&gt;
puts it plainly: to build RAG, you first need to create a vector store. Then the shortlist writes itself: Pinecone, Qdrant, Weaviate, Chroma.&lt;/p&gt;

&lt;p&gt;The vectors live in Postgres instead, through pgvector.&lt;/p&gt;

&lt;p&gt;Not because Postgres is faster. Because a new datastore is a new failure surface, with its own consistency model, its own operational habits and its own way of going wrong at 2am. Postgres was already there holding the documents, and I already knew how to debug it. Scaling was never the plan for a corpus of one rulebook.&lt;/p&gt;

&lt;p&gt;The whole of the retrieval query is this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;distance&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;distance&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;&amp;lt;=&amp;gt;&lt;/code&gt; is cosine distance: 0 is identical, 2 is opposite. The &lt;code&gt;::vector&lt;/code&gt; cast is not decoration. A bare Python list arrives as &lt;code&gt;double precision[]&lt;/code&gt;, which that operator refuses, so the vector travels as a text literal and is cast on arrival.&lt;/p&gt;

&lt;p&gt;One operator, one cast, one index. That is the entire integration.&lt;/p&gt;

&lt;p&gt;Convenience is not proof, though. Would the real vector database have bought back time worth having?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0l3semumf4xjuvgep0z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0l3semumf4xjuvgep0z.png" alt="Where one question's 2,823 ms actually goes: generation 2,598 ms, embedding API 219 ms, vector search 4.9 ms" width="800" height="323"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One question, end to end, takes about &lt;strong&gt;2,823 ms&lt;/strong&gt;. The vector search inside it takes &lt;strong&gt;4.9 ms&lt;/strong&gt;. Replacing pgvector with something infinitely fast would return 0.17% of the wait.&lt;/p&gt;

&lt;p&gt;None of this says a vector database is wrong. It says the unfamiliar one would have been paid for in operations and returned a line nobody was waiting on.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. You can only measure what you own
&lt;/h2&gt;

&lt;p&gt;No LangChain, no LlamaIndex either. The pipeline is plain Python calling two LLM vendors and a Postgres.&lt;/p&gt;

&lt;p&gt;The usual defence is YAGNI, and it holds: the flow is five fixed steps with no branching. Embed the question, search, build a prompt, call the model, parse the answer and its citations. Orchestrating five fixed steps is the easy part. A framework earns its keep on branching, retries and swappable backends, and none of those were in play.&lt;/p&gt;

&lt;p&gt;But that is not the reason worth giving.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2b23ndaquxjrmj3rxypm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2b23ndaquxjrmj3rxypm.png" alt="Five fixed steps; the three I wrote are the three that produced findings" width="799" height="361"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every result worth publishing from this project came out of a layer a framework would have owned.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I wrote the chunker, so the repeated headers could be stripped from every page and the effect measured: retrieval did not move at all. I wrote the rank fusion, so it was visible that hybrid search scored exactly the same as dense-only, and that the corpus was too small to show a difference rather than the technique being wrong. I wrote the database call, so it could be timed alone and found to be 4.9 ms.&lt;/p&gt;

&lt;p&gt;Three findings, all negative, all only visible from inside.&lt;/p&gt;

&lt;p&gt;The cost is real and there is a file listing it: mature libraries exist for the chunker, the fusion, the metrics, the eval harness. That file is the off-ramp for when this stops being worth it.&lt;/p&gt;

&lt;p&gt;You cannot instrument an abstraction you did not build.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Two dependencies called "model", two very different bills
&lt;/h2&gt;

&lt;p&gt;The model this system was planned around was retired before the first commit. The decision record replacing it carries the same date as that commit. The project lost its LLM before it had a second file.&lt;/p&gt;

&lt;p&gt;The swap itself was trivial. The responder is one environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;LLM_MODEL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LLM_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-haiku-4-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line, one deploy, done. Now the other model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdqqwkgdhmm9aexi7nb0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdqqwkgdhmm9aexi7nb0.png" alt="Blast radius: swapping the responder touches one config value; swapping the embedding model touches every stored vector" width="800" height="1183"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The embedding model cannot move like that. The table declares its vectors as 1536 dimensions and an insert guard rejects anything else, because vectors from different embedding models do not share a space. Changing it means re-embedding the entire corpus before a single query works again. An afternoon at this size. A migration at a real one.&lt;/p&gt;

&lt;p&gt;Same word, "model". One is a config value. The other is infrastructure wearing a config value's clothes.&lt;/p&gt;

&lt;p&gt;The vendor's deprecation schedule does not read your roadmap. Decide which of your dependencies is which before it decides for you.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The localhost trap: a 10-second database connection on Windows</title>
      <dc:creator>Serhiy Kucherenko</dc:creator>
      <pubDate>Mon, 27 Jul 2026 15:41:36 +0000</pubDate>
      <link>https://dev.to/skucherenko/the-localhost-trap-a-10-second-database-connection-on-windows-3le7</link>
      <guid>https://dev.to/skucherenko/the-localhost-trap-a-10-second-database-connection-on-windows-3le7</guid>
      <description>&lt;p&gt;Every question in the app took about 15 seconds. The model was fast. Retrieval was fast. Ten of those seconds were hiding somewhere else, and for days I couldn't say where.&lt;/p&gt;

&lt;p&gt;The app is payments-rag, a RAG system I built over the SEPA payment rulebooks. You ask a question in plain English, it retrieves the relevant passages from Postgres + pgvector, sends them to Claude, and returns an answer with the exact rulebook page cited. Fast is not a goal; under 5 seconds is. On my Windows machine it was taking three times that, every single query, and the answers were correct, which somehow made it worse. Correct but slow doesn't scream "bug." It whispers "maybe that's just how it is."&lt;/p&gt;

&lt;p&gt;For days I told myself it was just slow on Windows. Docker overhead, antivirus, something environmental, whatever. I checked the model latency three times before I thought to check the connection. That was the wrong order, and it cost me the better part of a week.&lt;/p&gt;

&lt;p&gt;There's a reason this class of bug survives. An LLM app carries a built-in excuse: models are slow, everyone knows models are slow, so a 15-second answer doesn't trigger the same alarm a 15-second SQL query would. The LLM had a good alibi for the slowness. I believed it too early.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring instead of guessing
&lt;/h2&gt;

&lt;p&gt;Two cheap pieces of observability ended the mystery in minutes.&lt;/p&gt;

&lt;p&gt;The first was a per-stage timer. I split each request into connect, retrieval, and generation, and logged the breakdown on every answer. The line read:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;retrieval 2.1s · generation 3.4s · connect + overhead 10.2s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retrieval plus generation was about 5.5 seconds. The other ten were sitting around the database connection, not in the model, not in the vector search.&lt;/p&gt;

&lt;p&gt;The second was a health check: a small panel that pings each dependency and shows the round-trip time. For the database it read:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB reachable · 10137 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Unambiguous. The connection itself took 10.1 seconds. Not the query, not the embedding, the TCP connect. The bug had been there for days behind a vague "it's just slow." The moment we measured, it took one screenshot to locate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually happening
&lt;/h2&gt;

&lt;p&gt;The connection string said &lt;code&gt;localhost&lt;/code&gt;. That one word triggered a chain of events.&lt;/p&gt;

&lt;p&gt;On Windows, &lt;code&gt;localhost&lt;/code&gt; is dual-stack: it resolves to both IPv6 &lt;code&gt;::1&lt;/code&gt; and IPv4 &lt;code&gt;127.0.0.1&lt;/code&gt;, and the OS prefers IPv6. So the Postgres client first tried to connect to &lt;code&gt;::1:5433&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Nothing was listening there. The Postgres in this setup runs in Docker, and the published port was bound on IPv4 only. The IPv4 loopback answers; &lt;code&gt;::1&lt;/code&gt; does not.&lt;/p&gt;

&lt;p&gt;And because no &lt;code&gt;connect_timeout&lt;/code&gt; was set, the client sat on the dead IPv6 route until the OS gave up on its own schedule and retried over IPv4. That giving-up takes about 10 seconds. Then the IPv4 connection succeeds in milliseconds, the query runs fine, and the answer comes back correct. Every query paid the toll again.&lt;/p&gt;

&lt;p&gt;Three things had to line up to make it hurt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;localhost&lt;/code&gt; is dual-stack, and Windows tries IPv6 first.&lt;/li&gt;
&lt;li&gt;Docker published the container port on IPv4 only, so there was nothing on &lt;code&gt;::1&lt;/code&gt; to answer.&lt;/li&gt;
&lt;li&gt;No &lt;code&gt;connect_timeout&lt;/code&gt;, so instead of failing fast and loud, the client waited out the full OS-level fallback on every single connection.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Remove any one of the three and the problem disappears, which is exactly why it survived for days. On Linux and macOS colleagues' setups the same code was quick.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;One word. &lt;code&gt;localhost&lt;/code&gt; becomes &lt;code&gt;127.0.0.1&lt;/code&gt;, which skips DNS ambiguity entirely and goes straight to the IPv4 loopback that Docker actually publishes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# .env&lt;/span&gt;
&lt;span class="nv"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;postgresql://user:pass@127.0.0.1:5433/mydb
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And a timeout, so this class of bug can never hide as a silent hang again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;psycopg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;connect_timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A hang is worse than an error. An error tells you where it hurts; a hang just eats your latency budget and says nothing.&lt;/p&gt;

&lt;p&gt;The before and after:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;host&lt;/th&gt;
&lt;th&gt;connect latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;localhost&lt;/code&gt; (IPv6 detour)&lt;/td&gt;
&lt;td&gt;10,137 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;127.0.0.1&lt;/code&gt; (direct IPv4)&lt;/td&gt;
&lt;td&gt;27 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;About 375× faster, one word changed.&lt;/p&gt;

&lt;p&gt;One caveat so nobody over-applies this: the trap needs the specific combination above. Windows resolution order, a container port published on IPv4 only, no connect timeout. If your stack differs, your ten seconds are hiding somewhere else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I keep from this
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;localhost&lt;/code&gt; is not &lt;code&gt;127.0.0.1&lt;/code&gt; once IPv6 is in play. For a container whose port is mapped to IPv4, prefer the literal address; it behaves identically on Linux and macOS, where loopback is loopback anyway.&lt;/li&gt;
&lt;li&gt;Always set a connect timeout. Fail fast and loud.&lt;/li&gt;
&lt;li&gt;Cheap observability pays for itself. A per-stage timer and a one-line health check are maybe an hour of work combined, and they turned a multi-day "it's just slow" into a five-minute fix.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The per-stage timer and the health check both stayed in the project permanently. They've earned their place: the same health view now checks all five dependencies on demand and every 10 minutes, so the next silent hang, wherever it comes from, gets a number attached to it before it gets a story.&lt;/p&gt;




&lt;p&gt;Serhiy Kucherenko builds backend systems and LLM tooling. The project this story comes from, payments-rag, is open source: &lt;a href="https://github.com/KucherenkoSerhiy/payments-rag" rel="noopener noreferrer"&gt;github.com/KucherenkoSerhiy/payments-rag&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>python</category>
    </item>
  </channel>
</rss>
