DEV Community

Cover image for The same question, asked three ways
Serhiy Kucherenko
Serhiy Kucherenko

Posted on

The same question, asked three ways

A RAG system answers a question in two steps. A retriever turns the question into a vector, compares it with the vectors of every passage in the documents, and hands the five closest passages to a language model. The model writes the answer from those five. If the passage that holds the answer is not among the five, the model cannot answer, however good it is.

In the previous piece I measured that second condition directly and found a question whose answer sat at rank 7. The model only ever saw ranks 1 to 5, so it refused. A reader, Mikhail, had found that question by hand and then found that two rewordings of it, both containing the words of the section that holds the rule, put the same passage at rank 1. He asked for a new kind of test question: the answer is in the documents, but the question does not use the section's own words. This piece builds that test and runs it.

The earlier experiment, in one paragraph

The documents are the two SEPA credit transfer rulebooks from the European Payments Council, 484 passages of about 300 words each, stored as vectors from OpenAI's text-embedding-3-small. The test set is 20 questions whose answer page I labelled by reading the PDFs, plus Mikhail's. The live system retrieves the five closest passages by cosine similarity, nothing else. Two variants exist in the code but are not switched on: collapsing byte-identical passages before taking the top five (138 of the 484 passages are exact duplicates, because a 36-page annex is bound into both rulebooks), and a hybrid path that also runs a Postgres full-text search and merges the two rankings.

flowchart LR
    Q["one question<br/>3 phrasings"] --> E["embedder<br/>text-embedding-3-small"]
    E --> R["ranking of all 484 passages<br/>by cosine"]
    R --> M["where is the labelled page?<br/>rank 1 .. 484"]
    Q -. "words only" .-> F["full-text search<br/>Postgres, english"]
    F --> R2["ranking by ts_rank"]
    R2 --> M
    R --> D["drop exact duplicates"] --> M

Note: the measurement is the rank of the labelled page in the full ranking, not a pass or fail from a judge. Rank 5 or better means the model would have seen the passage. Rank 6 or worse means it would not, whatever it then answered.

Hypothesis

Three phrasings per question:

  • original: the question as it stands in the golden set, written by someone who knows the rulebook, so it shares the section's vocabulary.
  • plain: the same need as a customer or a new joiner would type it. "Bank" instead of PSP, "instant transfer" instead of SCT Inst, and none of the terms the section is built on. Each question has a list of forbidden terms and the runner refuses to start if the plain phrasing contains one.
  • terse: three to five words, what a person types into a search box. It may reuse section terms; here the variable is length.
Original Plain Terse
How long does the Beneficiary PSP have to send a Return in a standard SCT? The account I paid to is closed. How many days does the receiving bank have to send the money back to my bank? closed account money sent back when
Does the SCT scheme itself set a maximum transaction amount? Is there a cap on how much I can send in one SEPA transfer? SEPA transfer cap
Within how many calendar days can a PSMB member request a telephone meeting instead of the written vote? (Mikhail's rewording) Within how many calendar days can a PSMB member request a telephone meeting? (his first question) PSMB telephone meeting days

What I expected:

  • H1. The plain wording pushes several answers out of the top five, and the terse wording more. Four questions already sat at rank 4 or 5 with their original wording, so at least those are exposed.
  • H2. Dropping duplicates helps the plain and terse phrasings more than the originals, because it frees slots that duplicate passages otherwise take.
  • H3. In the language piece the hybrid path cost 3 questions in 20 on English questions, and I blamed passages matching on a single shared acronym. A rule that only counts a passage when it matches more than one query word should recover most of that.
  • H4. A rule that switches the lexical leg off when the query language differs from the corpus changes nothing here by construction; every phrasing is English.

The first held. The second and third did not.

Development

Everything runs against the stored production vectors over a read-only connection. Each phrasing is embedded once, compared with all 484 passage vectors, and the rank of the first labelled page is recorded. The full-text leg is the production query: plainto_tsquery with its terms joined by OR, ranked by ts_rank. The hybrid leg merges the top 20 of each list with reciprocal rank fusion, as the code does. Two extra legs implement the H3 rule with a threshold of 2 and of 3 distinct matched query words.

Mikhail's question is the twenty-first row. Its original is his rewording that answered on the live demo, its plain is the question he actually asked first and was refused.

Note: I wrote the plain and terse phrasings before the run and read them back against the labelled pages. Golden sets are the ground truth here, so changing one after seeing the ranks is not allowed; these are the phrasings the run used.

Results

Phrasing alone, vector retrieval

Phrasing Answer in top 5 of 21 Mean reciprocal rank
original 0.81 17 0.61
plain 0.67 14 0.38
terse 0.43 9 0.27

The plain wording loses 3 of the 17 hits and gains none: the weekend-availability question falls from rank 2 to 46, the value-limits question from 4 to 8, and Mikhail's from 1 to 7. The terse wording loses 9 and gains 1. Eight of the 21 questions stay in the top five under all three phrasings, three are outside it under all three, and the remaining ten depend on the wording.

The plain wording is not uniformly worse, it is different. It moves the standard execution-time question from rank 5 to 2, the recall-deadlines question from 4 to 1, and the charging-principle question, a miss at 38 with the original wording, to 15. The embedder does not prefer the rulebook's vocabulary; it prefers whatever the passage happens to say, and the passage on Recall deadlines happens to say "Duplicate", "Fraudulent" and "sent".

The weekend question is the clearest loss. "Is the SCT Inst scheme available on weekends and public holidays?" finds the page that says "available 24 hours a day and on all Calendar Days of the year" at rank 2. "Can I send an instant transfer on a Sunday or on a bank holiday?" finds it at rank 46. The passage never says Sunday, weekend or holiday. It says "Calendar Days". A reader knows those are the same thing; the embedder puts them 44 places apart.

Dropping duplicates, on the full set

Phrasing vector duplicates dropped
original 17 17
plain 14 15
terse 9 10

One question flips, and it is Mikhail's, from rank 7 to 4 under both his first wording and the terse one. No other rank changes on the 21 originals, and every other move on the plain and terse phrasings is one to three places, all of them outside the top five already. The duplicate annex is 28.5% of the index, but it only crowds the top five when the question is about the annex, and none of the 20 golden questions is. H2 was wrong: dropping duplicates is the right fix for exactly one question in this set.

The hybrid path and the rule I promised

Leg original plain terse
vector 17 14 9
hybrid (production shape) 14 8 11
hybrid, passage must match 2 words 14 8 10
hybrid, passage must match 3 words 14 8 12

On the original questions the hybrid path loses three: Recall deadlines (rank 4 to 7), value limits (4 to 10), Return deadline (2 to 7). They are the same three the language piece lost, so that measurement reproduces. On the plain questions it loses six more, because "bank", "money" and "transfer" match passages all over both rulebooks. On the terse questions it gains two, since a three-word query is a lexical job and the full-text leg does it well.

The minimum-words rule does nothing to the three lost questions. The rows are identical. I had expected the noise to be passages that share one acronym with the question, and a threshold of two would have removed them. That is not what the full-text leg returns for an English question:

flowchart LR
    Q["Does the SCT scheme itself<br/>set a maximum transaction<br/>amount? (6 query words)"] --> FT["full-text top 20"]
    FT --> N1["rank 1: SCT Inst p78<br/>obligations list<br/>matches 6 of 6 words"]
    FT --> N2["ranks 2 to 10<br/>match 5 of 6 words"]
    FT --> N3["ranks 11 to 20<br/>match 4 of 6 words"]
    FT -.-> G["labelled page, SCT p21<br/>Value Limits<br/>rank 93, matches 3 of 6"]
    style G fill:#FBEAE3,stroke:#993C1D

Every passage in the full-text top 20 matches four to six of the six query words. The labelled page matches three and sits at rank 93. The same shape holds for the other two lost questions: the top 20 match four or five words each, the answer matches four and sits at rank 16 and 32. A threshold of two or three removes nothing, and the fusion goes on trusting a list whose top is wrong.

So the loss on English questions is not single-word noise. It is the lexical leg ranking the right passage poorly, because a long question shares many ordinary words ("scheme", "set", "transaction", "amount") with many passages, and the one passage that answers it is not the one that repeats the most of them. The acronym explanation from the language piece was right for foreign-language questions, where the acronyms were the only words that could match, and wrong for English ones. H3 fails, and the rule I promised in the comments does not touch the problem it was meant to fix.

Mikhail's question, all legs

Phrasing vector duplicates dropped full-text alone hybrid
his rewording (original) 1 1 1 1
his first question (plain) 7 4 1 3
terse 7 4 1 4

This row reproduces the previous piece to the digit. It is also the only row in the set where the full-text leg has the answer at rank 1 under every phrasing: his first question has 9 words and the passage that holds the rule contains 8 of them, "PSMB", "telephone" and "meeting" among them. The lexical leg rescued his question because his question is the kind the lexical leg is good at. It does not follow that it rescues the others, and on this set it does not.

Conclusion

  • Wording moves retrieval more than any model swap did. A plain rewording of the same need costs 3 answers in 21; a search-box fragment costs 8 net, 9 lost and 1 gained. The model matrix piece found 11 of 80 questions depending on the model choice; here 10 of 21 depend on how the question is put, with the embedder held fixed.
  • The loss is not "rulebook words good, plain words bad". Plain wording improved three ranks and lost three hits. What decides it is whether the words the user chose happen to appear in the passage: "Sunday" against "Calendar Days" loses, "twice by mistake" against "Duplicate sending" wins.
  • Dropping exact duplicates is a fix for one question in this set, the one that is about the duplicated annex. It costs nothing and it is still worth doing, but it is not the reason the golden set misses.
  • The hybrid path helps short queries and hurts long ones, and the minimum-matched-words rule does not change that, because on English questions the full-text leg ranks the right passage poorly rather than letting noise in. That corrects my own explanation from the language piece.
  • The test type Mikhail asked for is worth keeping: same question, no section words, scored by the rank of the labelled page. It caught the weekend question, which passed every previous test in this series because every previous test used the rulebook's own words.

Out of Scope

  • Whether the model answers correctly from the plain phrasing when the passage is retrieved. This piece stops at the ranking.
  • The language-match rule (H4), which needs the Spanish and Russian questions from the language piece. Not run.
  • A better lexical ranker (BM25 instead of ts_rank, or field weights) as the fix the hybrid path actually needs. Argued from three questions, not built.
  • More phrasings per question. Three is enough to show the effect and too few to give a per-question stability number.

Sources

  • The runner and the phrasing set, in the payments-rag repo: comparison/paraphrase/run.py, comparison/paraphrase/golden/sepa-phrasings.yaml.
  • Page labels: evals/retrieval_golden_set.yaml, verified against the PDFs.
  • The answer was at rank 7 (the refused question, the duplicate annex).
  • Asking a RAG in the wrong language (hybrid at a cost of 2 to 3 in 20, the acronym explanation this piece corrects).
  • Swapping every model in a RAG (11 of 80 questions model-dependent).
  • EPC SCT and SCT Inst Rulebooks 2025 (the corpus).

Top comments (0)