The cataloguing step is supposed to be the easy win. You have a schema whose
tables are called ecm_template_link and v_pmpm, your users ask questio...
For further actions, you may consider blocking this person and/or reporting abuse
Follow-up, and this one has an actual answer.
Yesterday I said the claim did not survive a proper test and that whatever was happening looked like it had measured the embedder. That was the honest reading of the numbers I had. It was not the cause. I found the cause.
The same column comment was being counted three times. It went into the prose channel, and into embed_text(), which feeds the body channel and the vectors. So a Spider 2.0 table carrying a description per column looked like a match on three of four channels for any question sharing a single word with any one of its columns, and the widest tables became magnets. Deleting the descriptions "helped" because it removed two thirds of a triple-count.
Decomposed on the same 212 questions: taking the comments out of the prose channel alone recovered 4 questions at k=10; out of embed_text() alone recovered 11. That is where the damage was.
The fix takes column comments out of the body and prose channels and leaves embed_text() alone, so the published vectors stay byte-identical and no stored index is invalidated. Re-run paired, same 212 questions:
Thirteen-to-one became two-to-two. The penalty is gone. Descriptions now neither help nor hurt on this benchmark, which is the floor a fielded design was supposed to guarantee and mine did not.
So the corrected claim is narrower and duller than the headline of this post, and it is the true one: long descriptions hurt retrieval in that implementation, not in general. The measurement was right. The inference was wrong. The right response to a true measurement is to fix the thing it measured, not publish it as a property of the world.
Separately, fixing the loader moved the resolvable set from 203 to 212 questions and k=20 from 67.2% to 71.3% on the 247-question denominator - while k=10 fell, 77.8% to 77.4%, because the nine questions that joined are harder than the average. Recording that rather than quoting only the cut that rose.
Shipped in 0.1.57. Grid and decomposition are in BENCHMARKS.md.
Recomputed the six cells of the before-table from (b, c) at
4b7b8ba1, and all six match to six decimals (16/9 → 0.229523, 19/5 → 0.006611, 15/3 → 0.007538, 13/1 → 0.001831, 12/3 → 0.035156, 15/2 → 0.002350), so "five of six significant" is exactly right — the k=5 hashed cell is the one that isn't. The after-table percentages check too. Three things the pair of tables can carry that they currently don't.The attribution names a function where the code names a channel. BENCHMARKS.md says the decomposition's arms were "the prose channel alone" (4 questions) and "
embed_text()alone" (11)._body_text's docstring says: "taking the comments out of this channel alone recovered eleven of the thirteen questions that deleting every description recovered (n=212, p=0.0225 at k=10)", and then "the vectors are deliberately not built from this.embed_textis a published, pinned guarantee". Those are two different interventions: removing the comments from whatembed_text()produces changes the body text and the vectors; removing them from the body channel changes one of its two consumers. The shipped fix does the second —texts = [embed_text()]still goes toself.embedder.embed(texts), so the vectors are untouched — and the penalty still disappears. Which means the 11 belong to the body consumer, not toembed_text()itself; if they belonged to the vectors, the fix as scoped could not have recovered them. That inference is the only reason I believe the exemption is free, and it's currently something a reader has to reverse-engineer from a docstring. Naming the channel in the BENCHMARKS sentence (or splitting the arm in two) makes it checkable.4 + 11 is 15, and the arm that removes every copy recovered 13. At k=10 with MiniLM the deletion arm is b=13, c=1. Two single-channel arms that sum to more than the arm containing both of them means they overlap by at least 2, or the third arm breaks something a partial removal left standing (c=1 there is consistent with either). The docstring's phrasing — "eleven of the thirteen" — is the union phrasing and is the correct one; the sum is an upper bound and reads as a decomposition. One table with b, c and p for all three arms plus the overlap would settle it, and it's cheap because you have the runs.
"The penalty is gone" is a claim about discordance, not about p. The after-table cells are d = 12 / 4 / 5 at k=5 / 10 / 20, and the two-sided exact floor
2^(1-d)is 0.00049 / 0.125 / 0.0625. The last two are above 0.05, so those cells could not have reached significance at any split, whatever the data said; p=1.0000 there is the absence of resolution rather than evidence of absence. Only k=5 could fire, and it returned 6/6. What actually carries the fix is a comparison the table doesn't show: the same arm, on the same 212 questions, before and after — b/c = 13/1 → 2/2 at k=10, i.e. d falls 14 → 4 (k=5: 25 → 12; k=20: 17 → 5). That is a within-arm change and it needs no test to be read; the resolution it buys is about 3.9 questions at k=10 (1.96·√d over 212), which is far below the 13 questions the penalty used to be, so the absence is meaningful because you can state what it excludes. Two columns (d, and the floor) would make that the visible reading.Two small ones. The after-table is the only table in the file that doesn't print the embedder, while the narrative follows MiniLM's cell throughout — worth labelling. And in the post's summary, "k=20 from 67.2% to 71.3% on the 247-question denominator" and "k=10 fell, 77.8% to 77.4%" are two different populations in one sentence: the k=20 pair is over 247, the k=10 pair is over the 203→212 comparison set. Over 247 the k=10 number must have risen, and it did — that's the 79.8% in the 247 block. The file is explicit about the convention ("both denominators are printed against every number here"); the sentence a reader quotes isn't, and someone recomputing it gets a contradiction.
One thing I'd keep from the post: "the measurement was right, the inference was wrong." Most withdrawals in this area are of the measurement, and that sentence is the difference between fixing the thing you measured and publishing it as a property of the world.
All five land. Three are now in BENCHMARKS.md, one is an edit to this post, and the first one is a documentation bug that was undercutting its own fix.
The attribution. You're right, and it's worse than a naming slip. The file said the 11 came from "
embed_text()alone" and then, two sentences later, that the fix leavesembed_text()alone. Both cannot be true: if the recovery had come from the vectors, a fix scoped to leave the vectors byte-identical could not have produced it._body_text's docstring had it right the whole time — "taking the comments out of this channel alone recovered eleven of the thirteen" — so the published prose was the wrong artifact, and you had to reverse-engineer a docstring to check the fix was sound. That isn't a reader's job. The sentence now names the body channel and says outright thatembed_text()is unchanged and still carries the comments.4 + 11 against a whole of 13. Not additive, and the file never said so. Each arm is measured against the same unmodified baseline, not against the other — removing one of three copies is sometimes enough to un-magnet a table that removing all three also fixes, so a question can come back under either intervention alone. Written down now.
The floor, which is the sharpest of the five. With d of 12, 4 and 5, the smallest attainable two-sided exact p is 2^(1-d) — 0.125 at k=10, 0.0625 at k=20. Those cells could not have reached p < 0.05 at any split whatsoever, and I left p=1.0000 sitting beside the words "the penalty is gone", where it reads as evidence. The table now tells you to read b and c and ignore the p column: thirteen-to-one became two-to-two, the discordance converged, and that is the only claim the rows carry.
The embedder. Hashed, and you're right that it was the one table in the file not saying so while the narrative around it follows MiniLM's cell. The identifying mark is now stated: the
without prosecolumn (143 / 177 / 187) is byte-identical to the hashed rows above, which is what you'd expect, since deleting every description is untouched by a change to which channel indexes them.The two denominators. This one is the post's fault, not the file's. BENCHMARKS.md does say "On the 212-question comparison set" for the 77.8 → 77.4; the summary here dropped it, so 67.2 → 71.3 (over 247) ends up in the same sentence as a pair over 212, and anyone recomputing gets a contradiction. Over 247 the k=10 number rose — 79.8% in that block, exactly as you say. I'm editing the post to carry both denominators rather than quietly fixing the number, since the mismatch is the thing worth seeing.
On "the measurement was right, the inference was wrong" — that's the one I want to keep too, and it's why the withdrawal is written into the file rather than replacing what it withdraws. A benchmark nobody can recompute is a press release.
Thank you for three days of this. Five for five, all real, all in a file I wrote and reread.
Verified all five against
f9516fe3: the attribution paragraph now names_body_textand says outright thatembed_text()is unchanged and still carries the comments; the non-additivity line (4 and 11 against a whole of 13) is in; the floor is in with 2^(1-d) and "read the b and c columns, not the p column"; the after-table printshashed embedder. The article itself still has no edit on it — dev.to reportsedited_at: nullfor 4641246 as I write this — so the two denominators are still only in the comment.One correction, and it is mine, not yours. "Thirteen-to-one became two-to-two": 13/1 is the MiniLM row of the before-table at k=10 and 2/2 is the hashed row of the after-table. The after-table has no MiniLM row at all, so that pair crosses the embedder axis, and the two cells cannot be one arm's before and after. My comment gave you that pair as "the same arm" — that was my error, and it is now the one sentence in the new text that does not survive its own tables. The same-embedder pairs, hashed, are:
Note the direction of my error: at k=10 I understated the drop (14 instead of 18), at k=20 I overstated it (17 instead of 15). k=5 came out right by accident — 25 is the hashed cell, 24 is MiniLM's.
The stronger form of "the penalty is gone" is in the levels, not the discordance, and it is same-embedder. Hashed, with prose: 136 -> 143 (k=5), 165 -> 177 (k=10), 178 -> 188 (k=20). The prose-removed arm read 143/177/187 before the fix and 143/177/187 after it. So the fix lands the with-prose column exactly on the prose-removed column at k=5 and k=10, and one question above it at k=20. Three cells matching exactly and one better is a stronger statement than b and c converging — it needs no test — and it states the residual, which is in the fix's favour rather than against it.
The new "one check that the table is the artifact it claims to be" is stated along the wrong axis. "As it must be, because deleting every description is untouched by a change to which channel carries them" is a reason the column is stable across the fix — and it would be stable across the fix under any embedder, so as written it cannot distinguish hashed from MiniLM. What gives it force is the column next to it: had the re-run been MiniLM, the without-prose column would read 154/176/189, not 143/177/187. Also "byte-identical to the hashed rows" is loose — those rows' with-prose cells are 136/165/178; the identity is with that arm's prose-removed cells. One clause fixes it: "identical to the hashed arm's prose-removed cells and not to MiniLM's."
Taking the correction, and it is the better version of the sentence. Landed in 6fea038.
You are right that the pair crossed the embedder axis. 13/1 is MiniLM at k=10; the hashed before-cell is 15/3, and the after-table is hashed throughout. So the sentence announcing that the fix worked contained the exact error the paragraph three lines above it warns readers about. Same embedder on both sides, hashed:
At k=10 the real collapse is larger than the sentence claimed, 18 -> 4 rather than 14 -> 4, and at k=20 it is smaller. I recomputed all three from the two tables in the file rather than from your comment, and they agree with you.
The levels being the stronger statement: agreed, and it is in the file now. Hashed, with prose: 136 -> 143 at k=5, 165 -> 177 at k=10, 178 -> 188 at k=20. At k=5 and k=10 the with-prose arm lands exactly on the prose-removed arm; at k=20 it passes it. Convergence of b and c says the penalty stopped. The levels say what it was worth, which is the thing a reader actually wants.
The check stated along the wrong axis: also right, and that one I should have caught when I wrote it. The interesting comparison is not hashed against MiniLM, it is before the fix against after it. The prose-removed arm is 143, 177, 187 on both sides, hashed throughout, because deleting every description cannot care which channel would have carried them. Only the with-prose arm was ever going to move, and it is the arm that moved. Reworded to say that.
Noted too that you corrected the direction of your own arithmetic in the same comment, unprompted, in both directions — understating at k=10 and overstating at k=20. That is the part of this thread I would point at if anyone asked what good review looks like.
Verified
6fea038against the two tables rather than against your comment: the arrows are right (25→12, 18→4, 15→5), the level column is right (136→143, 165→177, 178→188 against a prose-removed arm that reads 143/177/187 on both sides), and "at k=20 it passes it" is the direction the numbers carry. A paragraph that was arguing from the wrong axis now states the check that can fail, which is the part I would keep.One thing the fix introduced, worth a line while the diff is fresh:
_body_textis a second copy ofembed_text, identical line for line except forparts.extend(c.comment for c in self.columns if c.comment). I cloned at6fea038and grepped the tree — the name appears three times, all incatalog.py. So the relationship the docstring states ("everythingembed_textcarries except the per-column comments") is carried by that sentence plus the two functions agreeing today. Nothing fails if they stop agreeing.The asymmetry is what makes it worth making explicit.
embed_texthas a test asserting on it —test_catalog.py::test_view_definition_is_indexed, the view's definition being indexed — and_body_texthas none. The tested builder is the one that did not change. So the first component added toembed_text(a column type, a sampled value) reaches the vectors, the name channel and the prose channel, and silently skips the body channel; the only symptom would be a retrieval number moving in whichever experiment is running. That is the failure the paragraph above the table warns about, one level down: two texts meant to differ in exactly one component can come to differ in two, and the benchmark would attribute the difference to the comments.Cheap to make impossible: build the body text as a filter over the same parts list rather than a re-listing, or assert in a test that the two strings differ by exactly the comment part, joined by the same separator. Either turns the docstring into a check.
I ran your own eval on your own code, since the claim is checkable and you asked for exactly this check. Below:
schemagateatmain,tests/run_paraphrase_eval.py, the checked-intests/descriptions/*.json, your 52 questions, your TUNE/HELDOUT split, the hashing embedder. Your floor and your headline reproduce: identifiers alone 29/52 = 55.8% against the 56% in your docstring, shipped configuration 49/52 = 94.2% against your 92%.The cheap fix does nothing here. A commenter suggested keeping the generated text out of the IDF statistics. I built it two ways on one flat bag: (a) term frequencies and lengths from the enriched text, idf computed over the description-free text; (b) that plus the length prior's
len/avgalso taken from the description-free text. Both give 41/52 - the same questions as the flat bag, per-question identical. So at 27-51 objects the idf collapse is not what is costing you anything. Your 1,245-object schema is a different corpus and I don't have it; I tried scaling by replicating non-gold objects up to 1,065 and the arms still didn't separate, but there the copies inflate identifier overlap as well, so that test is confounded and I wouldn't read anything into it.The flat bag is not a regression on these schemas. 41/52 against 29/52 for identifiers alone. Concatenation does not reproduce the production failure here, which is consistent with the mechanism being scale-dependent: the damage needs a term to reach half the corpus, and at 42 objects a sentence per object can't get there.
What carries the gain is the prose as a peer channel. Nulling channels on top of your configuration: prose channel off 32/52 (below the flat bag's 41), union channel off 44/52, name channel off 44/52. So the descriptions are worth about +16 questions (32 -> 48), and rank fusion is what lets them pay without competing with identifiers for term frequency inside one bag. That is a different claim from "concatenation was the bug" - on this eval the fielded structure is a gain, not a floor.
The number for the fix's own claim, on your split. +6 questions on TUNE (14/22 -> 20/22) and +1 on the four held-out schemas (27/30 -> 28/30, the +1 being warehouse 7->8). Your harness holds those out on purpose, so that +1 is what the held-out evidence for the fielding claim comes to on these fixtures.
One thing about the metric.
select(q, top_k=6)returns 7-16 tables, not 6 - the coverage additions are appended past the cut (commerce 9-13 items, warehouse up to 16). So the harness number is "gold anywhere in a 7-16-item list". Slicing to the actual first six costs the floor 2 questions (29 -> 27) and leaves the other arms within one, so this is a naming issue rather than the effect - but it does mean the appended family members are being counted as hits.And you asked what makes a corpus immune. The document-frequency check you recommend asks about the words users type. On this eval that check cannot fire: the paraphrase questions are built to share no tokens with the corpus, so their df is 0 before and after enrichment - I ran it across all five question sets and every content word is 0 -> 0. That is the honest shape of the answer: the regression you hit needs queries that do share vocabulary with the corpus, because the lexical channel is where pollution shows; while the gain on this eval arrives through prose-as-channel and the vectors, where no df check can see anything. Two modes that don't overlap. A corpus is immune to the one you describe exactly when its questions avoid the corpus's own vocabulary - and maximally exposed when they don't.
A bug, small:
run_paraphrase_eval.pybuilds 5 of 6 schemas for me.complexraisessqlite3.OperationalError: unknown database "billing"out ofcon.executescript(m.DDL)inbuild(), so its numbers - and the held-out aggregate that includes it - are missing when the runner is used as checked in.Two questions. (1) Does the probe above change anything on your 1,245-object schema - one flat bag, idf from the description-free text? If it does there, the mechanism is real and my "does nothing" is a size effect; if it doesn't, fielding is doing more work than pollution removal. (2) For the queries that degraded in production, do they share tokens with the table text (i.e. df-checkable), or are they disjoint like these paraphrase questions? That decides whether your diagnostic can see the next occurrence of this.
This is the most useful thing anyone has done with my code and several parts of it are corrections I have to accept.
The top_k finding is worse than you found, and it's mine to fix. You caught it on the paraphrase eval; it also applies to my public benchmark numbers.
select()fills to top_k, then appends up to 3 coverage objects, then FK expansion adds more — the code comment literally says "only ever additive". Andbenchmarks/spider.pyandbenchmarks/spider2.pyboth scoresel.table_nameswith no slice. So every figure I've published labelled "top_k=5" is really "gold anywhere in the set returned when asked for 5", which may be nine or twelve. The numbers aren't invented but the column heading is wrong. I'm going to report requested top_k, the median and p90 of objects actually returned, and recall — because the library really does send all of them, so slicing would understate the token cost while the current label overstates the precision. Either way it needed saying and you said it first.On the flat bag not being a regression at 27-51 objects: accepted, and it sharpens the claim rather than killing it. The mechanism needs a term to reach a large fraction of the corpus, and a sentence per object across 42 objects cannot get there. Which means my own shipped fixtures cannot demonstrate my headline finding — worth me stating plainly rather than leaving for a reader to discover.
On "fielded structure is a gain, not a floor": you're right that these are different claims and I've been sliding between them. On your eval the prose channel is worth +16 questions and fielding is what lets it pay; on the 1,245-object schema fielding is containment. Both can hold — it depends whether prose is a net contributor or a net polluter at that scale — but I've been writing the containment story as though it were the whole thing.
The df diagnostic point is the sharpest thing here. You're right that it cannot fire on paraphrase questions that share no tokens with the corpus, and that the two failure modes don't overlap. "A corpus is immune exactly when its questions avoid the corpus's own vocabulary, and maximally exposed when they don't" is a better statement of scope than anything in my post.
Your two questions:
(1) I can run the flat-bag-with-clean-idf probe on the 1,245-object schema — it's private so I can't hand it over, but I can report the number, and I'll say which way it goes.
(2) Yes, they share tokens. The failing production queries were things like "show the contacts of X" against a table literally named
contacts, which is precisely the df-checkable case. That's consistent with your framing and it means my diagnostic is scoped to the mode I actually hit, not to yours.Filing the
complexschema crash —unknown database "billing"out ofexecutescript— as a bug. Held-out aggregate being silently short one schema is the kind of thing that quietly flatters a number.Thank you for doing this properly. It cost you real time and it's improved the work.
Ran the numbers you said you'd produce, on the public fixtures at your commit
9c68ec6f8(2026-09-13T22:30Z, i.e. 45 minutes before your reply — which makes one sentence of that reply one commit stale).The coverage half is already fixed. In that commit
select()enforces the budget for coverage picks ("Inside the budget, not beyond it: drop the weakest ranked pick rather than grow the answer"). Measured over the 52 questions oftests/run_paraphrase_eval.py: ranked picks are exactly 6 per query (hybrid 166 + vector 146 = 312 = 6.0 × 52) andcoversfires zero times. So "fills to top_k, then appends up to 3 coverage objects" no longer describes HEAD.The remaining growth is 100% join closure. With the default
expand_fks=True, on those same 52 questions: returned median 10, p90 12, max 16, and 52/52 queries return more than 6 objects; mean growth +4.13; every one of the 215 extra objects carriesreason="fk". The returned set is not "top-k with slack", it istop-k ranked ∪ FK closure, and the closure is ~41% of the answer on these schemas. That matters for the fix, because the two halves need different decisions: ranked picks are a budget, FK closure is set completion, and no single column heading can carry both.What the mislabel costs, both readouts. As published (any gold in the full returned set, FK on): 29/52 = 55.8%. Restricted to the first 6 — what "top_k=6" claims: 27/52 = 51.9%, and the two questions are
warehouse(5→4) andtelemetry(6→5), both held-out. Making the label true by construction (expand_fks=False) also gives 27/52, with exactly 6 objects everywhere. So the mislabel costs 3.8pp of recall, all on the held-out side, and your proposed fix (requested k + median/p90 returned + recall) keeps the stronger configuration while making the heading honest. I'd state it that way rather than as a retraction: what the label inflated is precision — the object count and therefore the token cost — not the recall. "Every figure labelled top_k=5 is wrong" understates how much of your work survives, and overstating a correction is its own error.Per-reason reporting is a one-liner.
Selection.hitsalready carriesScored.reason, soCounter(s.reason for s in sel.hits)gives requested-vs-delivered plus its cause for free. One small thing while you are there: thereasondocstring inmodels.py:261still enumerates five values (hybrid | vector | lexical | fk | pinned) whilecatalog.pywrites a sixth,covers— a list that is a second copy of a fact the code owns, with nothing keeping the two in sync. That is the same shape as the thing this whole exchange has been about.Two things still open. The
complexfixture still throwssqlite3.OperationalError: unknown database "billing"at this HEAD — I re-ran it just now — so the held-out aggregate is silently five schemas out of six, which is the direction that flatters. And your df diagnostic: on these 52 questions only 33% contain a single token with df>0 in the name index at all (commerce 25%, health 10%, warehouse 30%, finance 40%, telemetry 60%). The other 67% are outside its reach by construction, so your scoping sentence has a number attached to it: on this eval the check could have fired on a third of the questions, and your production queries were in that third.Thanks for taking the top_k point past where I took it — the
spider.py/spider2.pyextension is the part I would not have found without your reply, and documenting that the shipped fixtures cannot demonstrate the headline is worth more than the headline.Right on every count. I reproduced it rather than take your word for it, and from the wheel rather than the tree —
pip install schemagate==0.1.52, the artefact that went to PyPI at 22:53Z and the one anyone actually gets:Identical to yours,
coversfiring zero times included.The stale sentence is staler than you were kind enough to say. The swap — "drop the weakest ranked pick rather than grow the answer" — landed in 5199001 on 12 Sep, a day and a half before I wrote that reply, not one commit. I described the coverage pass from the comment at
catalog.py:921, which still reads "Bounded, and only ever additive -- it cannot displace a ranked pick", and never read the forty lines under it that displace a ranked pick. So I committed yourmodels.py:261defect myself, in the same file, about my own code, in a reply whose whole point was that I had audited it. Both comments ship in 0.1.52.And the correction was overstated. Taking that too. "Every figure labelled top_k=5 is wrong" is the wrong shape: the returned set is top-k ranked ∪ FK closure, the closure is ~41% of the answer on these schemas, and what the heading inflated is object count and token cost — not recall. 3.8pp, and buying the honest label with
expand_fks=Falsecosts exactly the two questions that make it up. So the label moves and the behaviour stays: requested k, median/p90 delivered, per-reason counts, recall.One thing in your data I hadn't seen.
coversfires zero times across 52 questions and five schemas. It has the longest comment in the file, it is the mechanism I have written about most, and the shipped eval never once exercises it. That is your paraphrase-eval finding again in a second place: my fixtures cannot demonstrate my claims. I would rather ship that sentence than the headline.Four filed: the unbounded FK closure, the two drifted comments, the
complexfixture (stillunknown database "billing"at HEAD — reproduced), and the reporting change.The 33% is the number I'll carry. A diagnostic that can fire on a third of questions is worth having and worth labelling as such, and commerce 25 / health 10 / warehouse 30 / finance 40 / telemetry 60 describes my fixtures' spread better than anything I have published about them.
Reproducing from the wheel rather than the tree is the stronger test and I should have done it that way round — 0.1.52 is what a reader actually gets, and your numbers match mine to the object count.
On
coversfiring zero times: I instrumented your guard on all 52 questions and it is narrower than "unexercised". It is unreachable.The admission test needs a token the name index knows (
idf >= 2.0) that the chosen names don't carry. Replaying that guard against the picksselect()actually returns: across the 52 questions, 223 tokens are absent from the name index entirely (idf 0) — all 52 questions contain at least one — and not one question has an admissible token that an already-chosen name doesn't carry, even with the idf floor lowered to 0. Every corpus-known word these questions use is already inside a ranked pick's name; the words that aren't are the ones the corpus has never seen, and the floor rejects them. The two sets are complementary on this eval by construction, because the high-idf name tokens are the same signal that drives the ranking that puts their objects in the picks — so the uncovered case only appears when a name-bearing token gets crowded out of a small budget.Which the sweep confirms. Running it at
expand_fks=True:Three firings in 364 question-runs, all of them at K <= 2, none at K >= 3, and the default is 6. So the pass you have written about most, with the longest comment in the file, and with the displacement behaviour you shipped on the 12th, has a reachable set that this eval samples zero times at every budget it reports numbers for.
Same sweep, second thing: the returned set exceeds the request at every K, not just at 5 — at K=1 it delivers mean 2.38 objects and up to 5, since a single ranked pick drags its FK closure in. "Requested k vs delivered" diverges from k=1 upward, so the reporting change is not a repair of one heading.
I looked for the four issues to link them from here and the public tracker shows six, newest #21 from 09-13T05:19Z, before our exchange — so they're still in your queue rather than on the record. Worth landing, because the numbers are the evidence and a later reader can't see a local list. The two stale comments are the same shape as the report you already fixed, in the same file, one of them written by me about it — a comment is a copy of what the code did, and it stops being one the moment the code moves under it.
Which is also why the 33% is the number I'd keep as the scope line: the diagnostic can reach a third of these questions, while the coverage guard reaches none of them, and both are fixtures' spread rather than a claim about retrieval in general.
Unreachable is right, and it is worse than your sweep showed. I ran K=1 through 20 over the same 52 questions — 1,040 question-runs, three
coversfirings, two at K=1 and one at K=2, zero at every K from 3 to 20. So it is not that the eval fails to sample it at 6. The guard has never fired at any budget anyone would use, including every budget BENCHMARKS.md reports numbers for.Same sweep answers your second point past where you took it:
It diverges from K=1 exactly as you said, and the absolute gap widens instead of closing — 4.5 objects over at K=20, 30 delivered against 20 asked at the worst case. So requested-vs-delivered is not a footnote on one heading; it is a column that has to appear at every cut.
On the issues you are right and I have nothing for it. Six open, newest #21 at 09-13T05:19Z, all of them predating this exchange. The four are a list in a chat window, which is the same failure you are describing one level up: the evidence is not on the record, so a later reader gets the claim without the finding. Filing them today, the two stale comments included.
And the 33% is the right scope line. A diagnostic that reaches a third of these questions and a guard that reaches none of them is a fair description of what these fixtures can show — which, after this week, is the sentence I would rather write than the alternative.
Update, and it goes against this post.
I re-ran it properly and the headline claim does not survive.
What changed. The Spider 2.0 comparison set went from n=158 to n=203, after a long-path bug turned out to be making 2,868 of 7,892 schema files silently unreadable. I added a cap sweep (descriptions truncated at 0 / 40 / 200 / 1000 words and unlimited), and tested it as the paired design it actually is: McNemar on the discordant pairs, not a comparison of marginals.
Twenty-four cells. One clears significance: MiniLM, descriptions deleted, k=20 - 15 questions fixed against 1 broken, p=0.001. Score the identical questions at the identical cut with the hashed vectoriser instead and it is b=11, c=6, p=0.332.
A result that survives one embedder and not the other has measured the embedder. It is also one cell out of twenty-four, where a Bonferroni threshold would be 0.002.
The sweep also removed the fix I was about to recommend: truncating at 200 or 1,000 words is indistinguishable from leaving the prose alone (one to four discordant pairs, p=1.000). Only deleting it entirely moves anything, and that is the cell above.
So "long descriptions hurt retrieval" is a signal worth chasing, not a result. The net margins in this post were real; the inference I drew from them was not earned.
One more thing that complicates it further, and I would rather say it than bury it. In Spider 2.0's schema files,
descriptionis not a table description at all - it is a per-column list aligned tonested_column_names, entries sometimes null, never a string (150 of 150 sampled). My loader joined it into a paragraph and indexed it as prose. So the ablation was deleting something that was never what its name suggested. Filed as a documentation request: github.com/xlang-ai/Spider2/issues...The full McNemar grid is in BENCHMARKS.md in the repo.
To the two of you who reported the same effect in dense retrieval: it may well hold there. I am no longer claiming it holds here at the strength this post claimed.
The corpus is the finding, and it is worth one more step, because this thread has now had the same correction twice: a worklist that was smaller than it looked. Your code comment has it — a table whose path is 260 characters long reads as a table that does not exist, so a 39% shortfall reads as a smaller database. I re-derived the printed grid rather than the claim; I have no reading to offer on the retrieval question itself.
The 24 cells. Recomputed from
(b, c)at3165d05: all twenty-four match your printed values at three decimals (15/1 → 0.000519, 11/6 → 0.332306, and so on). Two things that sit next to the numbers rather than changing them.The one cell is not a squeak.
pfor b=15, c=1 is 34/65,536 = 0.000519. Against the threshold your own sentence sets (0.05/24 = 0.002083) it clears by 4.0×, not by a hair. The "0.001 squeaks under" line is a three-decimal artifact of the grid's formatting — the margin is twice what the printed value makes it look like. If the cell is to be discounted, let the embedder swap do it, which does that work on its own: it does not need the multiplicity sentence as well.The multiplicity sentence and the floor. For
ddiscordant pairs the smallest attainable two-sided exactpis2^(1-d)— 0.0625 at d=5. Twelve of your twenty-four cells have a floor above 0.05, so no split whatsoever could have produced p<0.05 in them: cap=40 at k=10 and k=20 hashed, everything at cap=200 with k>5, and all six cap=1000 cells. The twelve that could fire are the six at cap=0 (d = 16–23), both at cap=40 k=5 (d=14), the two MiniLM cells at cap=40 k=10/k=20 (d=7), and both at cap=200 k=5 (d=8, 9). "One cell of twenty-four" is read as "one chance in twenty-four"; the grid's owndcolumn says something narrower — one of twelve cells that could fire, and the half of the grid where the manipulation moved at most 9 of 203 questions is where the floors sit. I am not proposing a different divisor: the family is 24 because 24 is what you ran. I am saying the sentence and the floor belong on one line, the way "the held-out p=0.031 rests on six discordant pairs" already does in the rerank table.Where the retired grid was measured.
benchmarks/cap_sweep.pystill builds the prose it caps:Two consequences. Null entries enter the paragraph as the literal token
None(your owntests/test_spider2_loader.pyfixture is["Unique row id", None, "Billing city"]), so at cap=40 some of the forty words kept can be the word "None". And[:cap]is a prefix of the concatenation, not each description truncated atcap: cap=40 is approximately "the text of the first few columns in index order". The sweep's docstring says "only the number of words kept fromdescription" changes — that is the one thing the code does not do; what changes is how much of the join survives. It is also a mechanical reason only cap=0 moves anything: cap=0 is the only arm whose manipulation is defined on the field itself. Every other arm manipulates a string the dataset does not contain, which is what you found by sampling, one level further down.c81bd223efixedbenchmarks/spider2.py;cap_sweep.pywas last touched byd345b8778, the long-path fix, and still joins. So the grid that survived is produced by the file the fix names, not by the fixed file. If the sweep is ever re-run, moving it ontospider2.load_dbgivescapa referent — per column — and then the reporting unit can be the corpus's own: for each cap, how many columns' text survived, and how many of the retained tokens are the literalNone. That turns the x-axis into "the first N of M columns", which is a statement about the schema rather than about a string the loader built. (This is the same unit problem as the idf threshold against the description counts in the other repo — the cut stated in a unit the corpus does not have.)What I did, so you can weigh it: recomputed your printed grid from
(b, c)and read three files at3165d05. I did not download Spider 2.0-lite, so the null-token and columns-survived counts above are what the code will produce, not what I measured. And my own numbers earlier in this thread — 41/52, the +16 for the prose channel — are on the paraphrase fixtures, a different corpus, so they do not inherit the path bug; but they are the same kind of count, and they would fail the same way under a worklist that was wrong.Last thing, and it is the part worth keeping: writing the withdrawal into the file instead of deleting the paragraph is what makes the Limitations list the right place for the promise-form hole you just confirmed. A reader can see both the claim and its removal.
Both mechanisms reproduce, and I ran them rather than reading them -- which is the one thing your comment explicitly did not do, so here are the outputs.
The literal None token is real. raw_desc joins the per-column list with str() applied to each element, so a null inside the list becomes the four-character word None inside the paragraph:
So the outer
return str(v) if v else Noneis not where it happens; it is thejoin(str(x) for x in v)one line above. A whole-field null stays a null, and a null element becomes a token the index treats like any other word.The cap is a prefix of the concatenation, exactly as you said. Same document, sweeping cap:
cap=4 is not "four words of each description". It is the first four words of the joined paragraph, three of them from column one and the fourth a null. So the x-axis is closer to "the first column and a bit" than to anything per-column, and cap=0 is the only arm whose manipulation is defined on the field the dataset actually has. That is the mechanical reason only cap=0 could move a number, one level below the sampling argument.
On the current state: benchmarks/cap_sweep.py on main still contains both lines, so a re-run as-is reproduces the same x-axis. BENCHMARKS.md does now carry the withdrawal -- "The cap sweep below was measured with the old loader and has not been re-run", with the cells kept "only as the record of what was done" -- which is the right call. What the note does not say is why, and the why is two checkable lines rather than a judgement: the cap applies to a join, and nulls enter it as a word. Someone re-running it later will otherwise reasonably assume the loader was the only problem.
Scope, so it can be weighed: I ran raw_desc and the cap line from current main against constructed documents. I did not download Spider 2.0-lite either, so I have no columns-survived or null-token counts for the real corpus -- only the demonstration that the code produces the shapes you predicted from reading it.
The collapse from 0.15 IDF explains so many mysterious retrieval regressions after document expansion passes. When an enrichment prompt summarizes a schema or a document cluster, the model naturally draws from the shared domain vocabulary. You end up inadvertently turning the most informative search terms into corpus-wide stopwords.
Separating fielded scoring makes sense here. For dense vectors, a similar inflation happens when generated descriptions compress every table into the same narrow semantic subspace. Keeping the raw identifiers in an isolated sparse channel is usually the only thing that keeps the exact table names from getting drowned out by the prose.
"Inadvertently turning the most informative search terms into corpus-wide stopwords" is a better one-line version than anything in my post.
The asymmetry I've landed on: expansion helps when the generated text adds vocabulary the corpus doesn't have, and hurts when it adds vocabulary the corpus already shares with the queries. doc2query over prose documents is mostly the first. Describing a schema is mostly the second, because the describing model has nothing to draw on except the domain's own nouns — it is, definitionally, writing the words your users type.
Your point about keeping raw identifiers in an isolated sparse channel is exactly where I ended up, and I'd now put it more strongly than the post does: it isn't only a hedge against bad prose, it's what stops good prose from doing damage.
Two things since I wrote this that sharpen it. On Spider 2.0-lite, which ships real data-dictionary descriptions rather than generated ones, deleting the descriptions entirely improved recall at every cut for both a hashed vectoriser and a sentence model. Those descriptions have a median of 155 words and a 90th percentile of 7,939 — documentation pages, not summaries. So the effect isn't specific to LLM-written text; it's about prose that shares vocabulary with the queries landing in the same field as the identifiers.
And I had a hypothesis that prose coverage explained why a sentence-transformer helped on Spider 1.0 (no descriptions) and not on Spider 2.0 (descriptions). I tested it by stripping the prose, and the embedder's advantage did not reappear. So that part was wrong — the encoder difference is something else, dialect or question phrasing or table size. Still working out which.
Correction, 13 Sep. The Spider 2.0 paragraph above does not stand up, for two reasons, and the second is the embarrassing one.
McNemar's exact test on the paired outcomes clears nothing — best cell p=0.065 (k=20: 9 fixed, 2 broken), and truncating descriptions instead of deleting them is p=1.000 at caps of 200 and 1,000, on one to four discordant pairs. "Deleting them improved recall at every cut" is a set of small net margins, which cannot tell you whether seven questions flipped one way or twenty-one flipped both.
Then the harness. It reads each table's JSON with a bare
except Exception: continue, and 3,056 of Spider 2.0's 7,892 table files have paths longer than the 260 characters Windows will open. So the ablation ran against a schema missing 39% of its tables, and the 155-word median and 7,939 p90 came from the same reader — floors, not counts. I had already found and fixed that bug in the scoring script; I did not fix it in the ablation script, and quoted the ablation anyway.So "the effect isn't specific to LLM-written text" is not yet supported by this evidence. The part above it that I still stand behind is the smaller claim: separating the fields is what bounds the damage either way, and that doesn't depend on the ablation.
The compounding is the nasty part: IDF collapse flattens the ranking, then length normalisation sorts what's left by inverse centrality, so the most central table gets hit twice. It also explains why down-weighting felt like the right lever but wasn't. IDF is a property of the bag, and the bag was already polluted. Scoring the fields separately gives each its own IDF space, which is what lets the description stay strong enough to bridge 'per member per month cost' to v_pmpm without dragging contact down to 0.15 in the name channel. What did you use to fuse the rankings, RRF or a weighted score sum? I've seen RRF hold up better when two rankings disagree hard, which is exactly the case here.
RRF, and
_RRF_K = 60. It fuses four ranked lists rather than two — vector, body-BM25, name-BM25, prose-BM25 — each contributingweight / (60 + rank + 1), summed.Your reasoning is the reason. A weighted score sum needs the scales to be comparable and they are not: BM25 over a three-token name and BM25 over a 7,939-word description produce numbers that do not mean the same thing, and every normalisation I tried smuggled a tuning constant back in through the side door. Rank is scale-free, so no single field's magnitude can poison the fusion — which is the same disease as the IDF collapse, one level up.
What RRF cost me, since you may hit it: it under-weights an exact name hit. A rare question word appearing literally in an object's name is the strongest signal available, and fusion flattens it to "rank 1 in one of four lists".
myconvoappeared in 5 names out of 1,245, matched the question exactly, and still sat seventh behind six tables that merely talked about conversations. So there is an explicit boost on top of the fusion for that case. RRF's scale-freeness is exactly what stops it recognising a certainty.And "length normalisation sorts what's left by inverse centrality" is a better sentence for the second half than anything in my post.
b=0.72was on the whole time — normalisation was working correctly and still made it worse, because the most central table is the one that attracts the longest description.Scale-freeness also hides the opposite of a certainty, a channel that knows nothing about the question. The floor argument in the post holds when a question shares no word with the generic prose, because every prose score is zero and the channel drops out. A question containing "users", or only "about", does share one. With "Stores data about users and their settings" on all 1,245 objects, that word has an idf of about 0.0004, still above zero, so every object gets the same positive score, the stable sort keeps them in the order they were indexed, and fusion pays the first one 1/61 and the fortieth 1/100 at the same weight as the name channel.
When the generic text varies only in length, the length prior breaks the tie instead, and the objects with the most prose about them sink, which is the second half of the original bug with a vote of its own. Putting one generic sentence in place of every description, shuffling the indexing order and rerunning the eval would show whether this ever moves a ranked pick. Giving tied objects one shared rank, or letting a channel abstain when its best query term has near-zero idf in that field, would make the floor true by construction.
You're right, and it is in the code rather than in the argument.
_rank()builds each channel's ranks like this:s > 0, then a stable sort, thenenumerate. So a channel that assigns every object the same positive score still hands out distinct ranks, and the order it hands them out in is the order objects were indexed.Your example, run on the 42-object demo schema with
"Stores data about users and their settings"as the description of every object, question"about users":A channel that knows nothing about the question pays 1.67x more to whichever object happens to sort first, at the same weight as the name channel. Exactly as you described it.
Then the part you asked about — whether it ever moves a ranked pick. It does. Permuting the index order eight ways and re-running:
So the tie is not cosmetic, and the baseline is not clean either: one question flips on index order even with real descriptions, which I did not know and would not have looked for.
One process note, since it belongs in this thread. My first attempt at that experiment shuffled the
CREATE TABLEstatements and reported 0 of 6, no effect. Reflection sorts table names alphabetically, so_orderwas byte-identical across all eight runs — the test had power zero and I was about to post its result as a refutation of your comment. Same shape as the 0/52 earlier in this thread. I only caught it because I printed_orderto check, and I now print it by default.Both of your fixes look right to me and they are not equivalent. One shared rank for tied objects is the local repair; abstention when the best query term has near-zero idf in that field is the one that makes the floor true by construction, because it also removes the length-prior tie-break you describe — which is the same "most prose about it sinks" mechanism as the original bug, with its own vote. I'd rather have the second and take the cost of deciding what "near-zero" means than have a floor that holds only when a question happens to share no vocabulary with the generic text.
Filing it with the reproduction above.
The zero-power run is the most useful thing in this thread, because it is the failure every shuffle test has by default: if the thing being permuted is re-sorted before it reaches the code under test, the experiment reports no effect with complete confidence. Printing _order catches it once; asserting it catches it every time, so the harness fails when two permutations produce the same order instead of relying on someone remembering to look. The one flip with real descriptions deserves the same treatment. It means some question still produces a tie on real prose, most likely on a word that is common across the schema, and listing the tied objects for that question would show whether abstention on near-zero idf removes it or whether it needs the shared rank as well.
Both fixes shipped, and the decomposition you asked for got run before they did.
_competition_rankgives equal scores the same rank and feeds all four channels -- vector, lexical, name, prose. That alone makes the fused score vectors identical across six indexing orders on 12 of 12 commerce questions in all three prose rows, against 3-4 of 12 that move without it. It does not close the output: 1-2 of 12 still change their top-6 while their scores are identical, because what survives is a tie at the final sort with nothing left to break it. That is the second anchor, one line away -- the key on the last sort,(-p[1], p[0])at catalog.py:991. 0 of 12 on both readings needs both.So to your question about the remaining flip on real prose: it is a tie, and abstention does not remove it. The guard fires only when the field's best query term is below 0.1 idf, which at N=1,201 means a term in roughly 1,088 of 1,201 documents. Real prose at these corpus sizes does not get there -- that is the whole reason the 1,200-object fixture had to be specified to produce it. Determinism closes that flip; the guard was never pointed at it.
Correction to my own paragraph, inside the hour. I first wrote that the assertion you are asking for is not in the tree. It is, and it is the version you named rather than the weaker one:
tests/test_tie_determinism.py.test_the_fixture_actually_tiesasserts the precondition before any conclusion is drawn from it -- every selected score equal -- so the suite fails if the fixture stops containing a tie, which is exactly the zero-power case.test_the_rank_does_not_depend_on_the_order_it_was_givenasserts a tie is present in the input before asserting stability across all its permutations.test_the_selection_does_not_depend_on_reflection_orderandtest_the_scores_do_not_depend_on_reflection_orderare your two readings, separately, plus a tie that straddles the top_k boundary. The module docstring says outright that both tests are written so that reverting either fix fails something. I asserted an absence from a directory listing I could not actually read, which is the same error this thread has been about from the start.What I find particularly interesting here is that the failure isn't really caused by “too much context”, it's caused by different types of evidence losing their distinct roles. A table name, schema relationship, and natural-language description answer different retrieval questions, so treating them as interchangeable text can make a strong exact signal look ordinary. This makes me think retrieval systems should preserve the semantics of each evidence source all the way through ranking, rather than only separating fields after the damage is done. The goal isn't necessarily less enrichment; it's making sure enrichment can't erase the signals that were already highly discriminative.
That is the mechanism, and there is a measured instance of it with a number on each side. RRF turns every rank into 1/(60+rank) and adds, so it rewards breadth over depth: an object first on two channels loses to objects placed tenth on three. On the 1,200-object fixture, "Top 5 contacts by email opens in the last 30 days, with their company" ranked
crm_engagement_factfirst of 1,199 on the body channel and first on prose -- its columns areemail_open_7d,email_open_30d-- and fusion put it 30th, behind twenty-ninecrm_contact_*andcrm_company_*siblings that merely share a word with the question in their name. Your "a strong exact signal made to look ordinary", exactly. The coverage pass could not help either, because it guarantees a slot for informative words in names, and "email" and "open" appear only in columns.The fix is structural rather than a weight, which is your point: the body channel is the only one that sees columns, so its single best hit is the one piece of evidence nothing else guarantees, and it now gets one slot under the same rules coverage uses -- budget-neutral, displacing the weakest ranked pick, never a pinned or covering one, abandoned rather than break the budget. Fielded recall at k=6 on that fixture goes 12/16 to 14/16 (fixes 3, breaks 1). On the six shipped schemas it is byte-identical, because there the best lexical match already made the cut.
Where I would resist the framing slightly is "only separating fields after the damage is done", because part of what I attributed to fielding turned out to be a bug propping it up.
_stemstripped "-es" unconditionally, soinvoicesnever metinvoiceand six of eleven common plurals never met their singular -- one side of a lexical index could not match the ordinary way anyone phrases a question. With that fixed the flat bag goes from 9/16 to 12/16 at k=6 and fielded scoring stays at 12/16, so the comparison I published no longer reproduces at the operating point. Foreign-key expansion on the same fixture went from worth 3 of 8 complex questions to 0 of 8, for the same reason.And the enrichment claim itself is now narrower than the post. McNemar on the ablation's discordant pairs clears significance in one cell of twenty-four, and that cell evaporates when the embedder is swapped: b=15 c=1 p=0.001 with MiniLM against b=11 c=6 p=0.332 hashed, same 203 questions, same cap, same cut. Preserving each evidence source's role is still what I would build. "Less enrichment" is not a result I can hand you.
That makes the distinction much clearer. I especially like the narrowing of the enrichment claim, because it shows why retrieval experiments need to separate the retrieval design from the quality of the underlying lexical pipeline. The stemming bug is a good example of how an apparently structural retrieval problem can actually be amplified by a lower-level implementation detail. I still think preserving evidence-source roles is useful as a design principle, but the updated results make the stronger point: each retrieval change needs to be tested against the full pipeline before attributing the effect to one component. That makes the ablation results much more actionable.
"Each retrieval change needs to be tested against the full pipeline before attributing the effect to one component" is the rule I wish I had written down first -- it would have saved me a fortnight. The stemming bug is the cheap version of the lesson: a lexical defect two layers below the thing I was studying produced a difference that looked structural, and I spent the time explaining the wrong mechanism.
Your design principle survives all of it and I would still build that way. What the ablation took away was not "give each evidence source its own role" -- it was my claim to have measured how much that is worth. Those are two different sentences, and I had been using the second to prop up the first.
Dense vector retrieval has the exact same problem and it's harder to catch because there's no IDF value to look at - you just see a ranking that feels slightly off and can't explain why. When I built a retrieval layer over a graph schema with about 400 nodes, enriching everything with generated descriptions created this clustering effect where all the high-centrality nodes ended up semantically adjacent to each other, so similarity queries returned a blob of related-but-wrong results instead of the specific node the question was about. Your fix of scoring name vs description separately before fusion is basically the same thing I ended up doing after two weeks of tuning field weights and getting nowhere.
The "no IDF value to look at" part is the bit I keep coming back to. In sparse you get a number telling you the term is worthless. In dense the same collapse just shows up as results that feel vaguely off, and nothing in the pipeline is red.
Your clustering description sounds like the dense form of exactly this: the high-centrality nodes attract the most description text, the descriptions share vocabulary, and they end up adjacent to each other rather than to the query. The sparse version has the same bias from the other side — the most central table has the longest document and gets punished hardest by length normalisation.
If it's useful as a diagnostic: mean pairwise cosine across the corpus before and after enrichment. If it goes up, you've compressed the space rather than described it. That's the dense analogue of watching document frequency, and I hadn't thought to suggest it until your comment.
Two weeks of tuning field weights before separating the fields is the same road I went down. Down-weighting is the intuitive lever and it's the wrong one, because the damage is done at index time — IDF is a property of the bag, so a lower weight scales a number that was already wrong.
One update since I wrote this, in case it saves you a step: I ran the same ablation on Spider 2.0-lite, which ships real data-dictionary descriptions rather than generated ones. Removing them improved recall at every cut, for both a hashed vectoriser and a sentence model. Their "descriptions" have a median of 155 words and a 90th percentile of 7,939. So it isn't only generated prose that does this — it's any prose that shares vocabulary with the queries and arrives in the same field.
Correction, 13 Sep. Withdrawing the Spider 2.0 paragraph above; two separate problems with it.
McNemar's exact test on the paired outcomes clears nothing — the best cell is p=0.065 (k=20, 9 questions fixed, 2 broken), and truncating descriptions rather than deleting them is p=1.000 at caps of 200 and 1,000. "Removed them and recall improved at every cut" is twelve small margins pointing one way, which is a direction to chase, not a result.
And the harness that produced them reads each table's JSON with a bare
except Exception: continue; 3,056 of Spider 2.0's 7,892 table files have paths over 260 characters, which Windows will not open. So it ran on a schema missing 39% of its tables. The 155-word median and 7,939 p90 come from that same reader, so they are floors rather than counts. I had fixed this in the scoring script and not in the ablation script.None of it touches the 1,245-object finding this post is about — different corpus, different harness — but I quoted it here as support and it isn't, yet.
This is the cleanest explanation I have seen of generated enrichment backfiring: you wrote 1,245 documents in one voice, and IDF correctly concluded that voice carries no information. The part that should worry anyone doing HyDE or generated summaries is that the descriptions were good. Accurate text can still be index poison. Keeping generated text in its own field with a lower boost, or excluding it from the IDF statistics entirely, is the cheap fix, but the deeper lesson is that the model's vocabulary is the bias, not the content. "Correlated, not wrong" is exactly the failure shape.
"You wrote 1,245 documents in one voice, and IDF correctly concluded that voice carries no information" — that's the sentence I spent a week failing to write. IDF didn't malfunction. It worked, on a corpus I had made uninformative.
On excluding generated text from the IDF statistics: that's subtly different from a separate field, and I think it only fixes half. You'd restore identifier IDF — "contact" goes back from 0.15 to 4.27 — but if the prose still sits inside the same document it still counts toward |d|, so length normalisation keeps punishing the 55-column central table that attracted the longest description. Separate fields fix both halves, because the name field's length is the length of the name and nothing can inflate it.
The lower-boost version I did try first, and it failed for the reason you'd expect once stated: the weight scales a contribution whose IDF was already computed over the polluted bag. It also cost me the queries only a description can answer, which is the whole point of having descriptions.
"Accurate text can still be index poison" is the part I'd want anyone doing HyDE to take away — quality control on the generated text doesn't help, because quality was never the failure mode.
Since posting, one thing that supports your framing harder than my own evidence did: I ran the ablation on Spider 2.0-lite, where the descriptions are human-written data-dictionary entries rather than LLM output. Removing them improved recall at every cut. Not a model's voice at all — just prose in the wrong field.
Correction, 13 Sep. Pulling the last paragraph. It was the strongest-sounding thing in this comment and it is the least supported.
McNemar's exact test on the paired outcomes clears nothing — best cell p=0.065 (k=20: 9 questions fixed, 2 broken), and truncating the descriptions rather than deleting them is p=1.000. "Removing them improved recall at every cut" describes twelve small net margins, not a measured effect.
And the ablation harness reads each table's JSON with a bare
except Exception: continue, while 3,056 of Spider 2.0's 7,892 table files have paths over the 260 characters Windows will open — so it ran on a schema missing 39% of its tables. I had fixed that in the scoring script and not in this one.So "human-written data-dictionary entries do it too" is not something I can claim yet. Which is a shame, because it was the part that supported your framing rather than mine — everything above it rests on the 1,245-object corpus and is unaffected.