The cataloguing step is supposed to be the easy win. You have a schema whose
tables are called ecm_template_link and v_pmpm, your users ask questions
in English, and the gap between those two vocabularies is why retrieval
misses. So you point a model at every table, get back a sentence describing
each one, index the sentences alongside the names, and now the corpus speaks
English too.
I did that to a real 1,245-object schema. Recall went down.
Not by a little. The table literally called contacts sat at rank 3 for the
question "show the contacts of xmagnet" before cataloguing. After
cataloguing it was below rank 40 — off the end of anything I would put in a
prompt. The descriptions were fine. I read them. They were accurate,
specific, and they made the system worse.
This post is what was actually happening, why the obvious fix doesn't work,
and the one that does. None of it is specific to text-to-SQL. If you are
enriching documents before indexing them — summaries, generated titles,
hypothetical questions, keyword expansion, anything — the same mechanism is
available to bite you, and it will not announce itself.
The mechanism, in one line
Every description you generate is written in the same vocabulary as every
other description, so enrichment raises the document frequency of exactly
the words your users type.
BM25 scores a term by inverse document frequency:
idf(t) = log(1 + (N - n_t + 0.5) / (n_t + 0.5))
N is the corpus size, n_t the number of documents containing the term.
A term in few documents is informative and scores high; a term in most
documents is worthless and scores near zero. That is the entire point of
IDF, and it is normally a good instinct.
Now think about what a model writes when you ask it to describe a table in a
CRM schema. It writes about contacts. It writes about contacts when
describing contacts, and also when describing contact_lists,
campaign_recipients, email_events, tenants, users, and the audit
table that logs changes to any of them — because in a CRM, almost everything
is about contacts in some defensible sense. The descriptions are not
wrong. They are correlated.
After cataloguing, the token contact appeared in roughly 1,072 of the
1,245 documents — a figure I can reconstruct from the IDF it produced,
which was 0.15.
For comparison, here is what the same schema gives you for a genuinely rare
term:
| term | documents containing it | idf |
|---|---|---|
contact, after cataloguing |
~1,072 of 1,245 | 0.15 |
contact, in the name field only |
17 of 1,245 | 4.27 |
tenant, in the name field only |
30 of 1,245 | 3.71 |
the |
0 of 1,245 | — |
The tokenizer does not strip stopwords, so the and and are in that
index too, sitting near zero because they are in everything. At 0.15,
contact had joined them. The word the user typed was, for scoring purposes,
a function word.
The second half, which is worse
IDF collapse alone would flatten the ranking. What actively inverted it was
length normalisation.
The b parameter in BM25 penalises long documents, on the sound theory that
a long document containing your term is less about your term than a short
one that contains it:
score += idf(t) * f * (k1 + 1) / (f + k1 * (1 - b + b * len / avg_len))
Ask which object in a CRM schema has the longest document, and the answer is
the central one. contacts in this schema has 55 columns. Add a generated
description and a row of alias words and its document is several times the
corpus average. Meanwhile contact_import_log has six columns and a
one-line description, so it is short, tidy, and — as far as the length
prior is concerned — much more about contacts.
So the two effects compound in the same direction:
- IDF collapse removes the signal that would have separated
contactsfrom the forty other tables mentioning contacts. - Length normalisation then actively sorts what's left by inverse centrality, because the most important table in a schema is reliably the one with the most columns.
Cataloguing didn't add noise. It added correlated noise, and correlated
noise attacks the exact query it was meant to help. The questions that
degraded most were the ones cataloguing exists to serve — plain English, no
schema words. Questions that named a table outright were mostly fine, because
they had a rare token to hang on. That is a nasty failure profile: the
feature looks fine on your smoke tests and fails on your users.
The fix that doesn't work
The obvious response is to trust descriptions less. One index, but weight the
generated text below the real text.
I did this first. It helps a bit and it is the wrong lever, for a reason
that took me a while to see: weight and dilution act at different stages.
Down-weighting scales the contribution of a term after IDF has already been
computed over a corpus the descriptions polluted. contact is still worth
0.15 in the name's own score, because name and description live in one bag of
words and IDF is a property of the bag. You have made a bad channel quieter
without making the good channel accurate again.
And the cost is real. Starving the prose weight cost me
"per member per month cost" → v_pmpm, which is the single best example in
the whole schema of a question only a description can answer. There is no
lexical path from that phrase to that name. The description was the only
bridge and I had just defunded it.
So: down-weighting trades away the wins to partially mitigate the losses. You
end up tuning a scalar that makes both worse than they need to be.
The fix that works
Score the fields separately and fuse the rankings, rather than concatenating
the fields and scoring once.
Three BM25 indexes over the same objects:
self._bm25 = _BM25([doc.embed_text() for doc in docs]) # everything
self._bm25_name = _BM25([_name_text(doc) for doc in docs]) # identifiers
self._bm25_prose = _BM25([_prose_text(doc) for doc in docs]) # written text
where
def _name_text(doc):
"""Just the identifiers: schema, name, and the name split on underscores."""
return " ".join(x for x in (doc.schema, doc.name,
doc.name.replace("_", " ")) if x)
def _prose_text(doc):
"""Everything written *about* the object: hint, description, comments."""
parts = [doc.hint or "", doc.description or ""]
parts.extend(c.comment or "" for c in doc.columns)
return " ".join(x for x in parts if x)
Then fuse by reciprocal rank rather than by score:
for q in candidates:
s = 0.0
if q in vec_rank: s += vector_weight / (RRF_K + vec_rank[q] + 1)
if q in lex_rank: s += lexical_weight / (RRF_K + lex_rank[q] + 1)
if q in name_rank: s += NAME_WEIGHT / (RRF_K + name_rank[q] + 1)
if q in prose_rank: s += PROSE_WEIGHT / (RRF_K + prose_rank[q] + 1)
Both halves of the bug die at once, and it is worth being precise about why,
because "just use fielded search" is advice people give without the
mechanism:
IDF is recomputed per field. In the name index, the only text is
identifiers. Nothing a model writes can ever enter it. contact appears in
17 names out of 1,245, so its IDF is 4.27 instead of 0.15 — 28× the
discriminating power, restored by construction rather than by tuning.
Length is per field too. The name index's document length is the length
of the name. contacts is two tokens whatever else you attach to the object.
The 55 columns cannot inflate it, so the length prior stops punishing
centrality.
Fusion is over ranks, not scores. This is the part that contains a bad
catalogue, and it's why I could raise the prose weight back to parity. A
channel can only ever contribute its own ranking. If a weak model writes
"Stores data about users and their settings" about all 1,245 objects, the
prose channel becomes uniformly useless — every object ranks the same, the
channel contributes nothing that discriminates, and the name and body
channels decide the result unchanged. The floor becomes "no better than
before cataloguing" instead of "worse than before cataloguing".
That last property is the one I actually care about. It means pointing a
small local model at your schema is safe. Not good, necessarily — a 1.5B
model writes considerably worse descriptions than a frontier model, and I'd
rather you use the good one. But safe: bad prose can no longer bury the
object it describes, so the downside of trying is bounded.
What this generalises to
The pattern is not about databases. It is:
Generated text about a corpus is written in the corpus's own vocabulary,
so enrichment inflates document frequency for the domain's central terms —
the ones users search with — and inflates document length most for the
items that matter most.
Anywhere you generate text and index it next to original text, in the same
field, you have signed up for both effects:
- Summaries prepended to chunks. Every summary in a corpus about Kubernetes says "Kubernetes".
- Hypothetical-question generation (HyDE-style indexing). You are synthesising the user's own phrasing, at scale, across every document. That is IDF dilution as a product feature.
- Keyword and synonym expansion. Same shape, more concentrated.
- LLM-written titles or alt-text merged into the body field.
None of these are bad ideas. I still catalogue schemas; recall on
business-phrased questions is far better with descriptions than without. The
claim is narrower: enrichment belongs in its own field, always. The cost
of separating fields is one more index and a fusion step. The cost of not
separating them is a regression that shows up only on your most important
queries and looks like "retrieval is just hard".
If you want to check whether this is happening to you, it is one query and
no instrumentation: take the ten nouns your users actually type, and print
their document frequency before and after your enrichment step. If any of
them are now in more than half your documents, that term is doing nothing,
and it was probably doing something before.
What it doesn't fix
Honesty about the edges, since the above reads tidier than the week did:
Fielded scoring does not make a bad catalogue good. It makes it harmless. If
your descriptions are generic, you get the pre-cataloguing ranking back, not
a better one — which is the right outcome, but don't read it as a licence to
skip evaluating the model that writes them.
It also introduces a knob per field, and I do not have a principled method
for setting them. Mine are all at parity because that tested best across six
schemas, not because parity is theoretically correct.
And separating fields cannot fix a term that is genuinely common in the
names too. A schema with 300 tables actually called contact_something has
a real ambiguity problem, and no amount of field isolation invents the
information to resolve it.
The measurements here come from schemagate, an open-source library
(Apache-2.0) that does the retrieval step for text-to-SQL. The relevant code
is in catalog.py
— the comments around the three _BM25 constructions are where I wrote this
down while it was still fresh. There's a browser demo at
ashishsinha1602.github.io/schemagate
that runs the real selector client-side on six sample schemas, if you'd
rather poke at the ranking than read about it.
If you run the document-frequency check on your own corpus, I'd like to know
what it says — particularly if it says nothing is wrong, because I'd like to
know what makes a corpus immune.
Top comments (81)
Follow-up, and this one has an actual answer.
Yesterday I said the claim did not survive a proper test and that whatever was happening looked like it had measured the embedder. That was the honest reading of the numbers I had. It was not the cause. I found the cause.
The same column comment was being counted three times. It went into the prose channel, and into embed_text(), which feeds the body channel and the vectors. So a Spider 2.0 table carrying a description per column looked like a match on three of four channels for any question sharing a single word with any one of its columns, and the widest tables became magnets. Deleting the descriptions "helped" because it removed two thirds of a triple-count.
Decomposed on the same 212 questions: taking the comments out of the prose channel alone recovered 4 questions at k=10; out of embed_text() alone recovered 11. That is where the damage was.
The fix takes column comments out of the body and prose channels and leaves embed_text() alone, so the published vectors stay byte-identical and no stored index is invalidated. Re-run paired, same 212 questions:
Thirteen-to-one became two-to-two. The penalty is gone. Descriptions now neither help nor hurt on this benchmark, which is the floor a fielded design was supposed to guarantee and mine did not.
So the corrected claim is narrower and duller than the headline of this post, and it is the true one: long descriptions hurt retrieval in that implementation, not in general. The measurement was right. The inference was wrong. The right response to a true measurement is to fix the thing it measured, not publish it as a property of the world.
Separately, fixing the loader moved the resolvable set from 203 to 212 questions and k=20 from 67.2% to 71.3% on the 247-question denominator - while k=10 fell, 77.8% to 77.4%, because the nine questions that joined are harder than the average. Recording that rather than quoting only the cut that rose.
Shipped in 0.1.57. Grid and decomposition are in BENCHMARKS.md.
Recomputed the six cells of the before-table from (b, c) at
4b7b8ba1, and all six match to six decimals (16/9 → 0.229523, 19/5 → 0.006611, 15/3 → 0.007538, 13/1 → 0.001831, 12/3 → 0.035156, 15/2 → 0.002350), so "five of six significant" is exactly right — the k=5 hashed cell is the one that isn't. The after-table percentages check too. Three things the pair of tables can carry that they currently don't.The attribution names a function where the code names a channel. BENCHMARKS.md says the decomposition's arms were "the prose channel alone" (4 questions) and "
embed_text()alone" (11)._body_text's docstring says: "taking the comments out of this channel alone recovered eleven of the thirteen questions that deleting every description recovered (n=212, p=0.0225 at k=10)", and then "the vectors are deliberately not built from this.embed_textis a published, pinned guarantee". Those are two different interventions: removing the comments from whatembed_text()produces changes the body text and the vectors; removing them from the body channel changes one of its two consumers. The shipped fix does the second —texts = [embed_text()]still goes toself.embedder.embed(texts), so the vectors are untouched — and the penalty still disappears. Which means the 11 belong to the body consumer, not toembed_text()itself; if they belonged to the vectors, the fix as scoped could not have recovered them. That inference is the only reason I believe the exemption is free, and it's currently something a reader has to reverse-engineer from a docstring. Naming the channel in the BENCHMARKS sentence (or splitting the arm in two) makes it checkable.4 + 11 is 15, and the arm that removes every copy recovered 13. At k=10 with MiniLM the deletion arm is b=13, c=1. Two single-channel arms that sum to more than the arm containing both of them means they overlap by at least 2, or the third arm breaks something a partial removal left standing (c=1 there is consistent with either). The docstring's phrasing — "eleven of the thirteen" — is the union phrasing and is the correct one; the sum is an upper bound and reads as a decomposition. One table with b, c and p for all three arms plus the overlap would settle it, and it's cheap because you have the runs.
"The penalty is gone" is a claim about discordance, not about p. The after-table cells are d = 12 / 4 / 5 at k=5 / 10 / 20, and the two-sided exact floor
2^(1-d)is 0.00049 / 0.125 / 0.0625. The last two are above 0.05, so those cells could not have reached significance at any split, whatever the data said; p=1.0000 there is the absence of resolution rather than evidence of absence. Only k=5 could fire, and it returned 6/6. What actually carries the fix is a comparison the table doesn't show: the same arm, on the same 212 questions, before and after — b/c = 13/1 → 2/2 at k=10, i.e. d falls 14 → 4 (k=5: 25 → 12; k=20: 17 → 5). That is a within-arm change and it needs no test to be read; the resolution it buys is about 3.9 questions at k=10 (1.96·√d over 212), which is far below the 13 questions the penalty used to be, so the absence is meaningful because you can state what it excludes. Two columns (d, and the floor) would make that the visible reading.Two small ones. The after-table is the only table in the file that doesn't print the embedder, while the narrative follows MiniLM's cell throughout — worth labelling. And in the post's summary, "k=20 from 67.2% to 71.3% on the 247-question denominator" and "k=10 fell, 77.8% to 77.4%" are two different populations in one sentence: the k=20 pair is over 247, the k=10 pair is over the 203→212 comparison set. Over 247 the k=10 number must have risen, and it did — that's the 79.8% in the 247 block. The file is explicit about the convention ("both denominators are printed against every number here"); the sentence a reader quotes isn't, and someone recomputing it gets a contradiction.
One thing I'd keep from the post: "the measurement was right, the inference was wrong." Most withdrawals in this area are of the measurement, and that sentence is the difference between fixing the thing you measured and publishing it as a property of the world.
All five land. Three are now in BENCHMARKS.md, one is an edit to this post, and the first one is a documentation bug that was undercutting its own fix.
The attribution. You're right, and it's worse than a naming slip. The file said the 11 came from "
embed_text()alone" and then, two sentences later, that the fix leavesembed_text()alone. Both cannot be true: if the recovery had come from the vectors, a fix scoped to leave the vectors byte-identical could not have produced it._body_text's docstring had it right the whole time — "taking the comments out of this channel alone recovered eleven of the thirteen" — so the published prose was the wrong artifact, and you had to reverse-engineer a docstring to check the fix was sound. That isn't a reader's job. The sentence now names the body channel and says outright thatembed_text()is unchanged and still carries the comments.4 + 11 against a whole of 13. Not additive, and the file never said so. Each arm is measured against the same unmodified baseline, not against the other — removing one of three copies is sometimes enough to un-magnet a table that removing all three also fixes, so a question can come back under either intervention alone. Written down now.
The floor, which is the sharpest of the five. With d of 12, 4 and 5, the smallest attainable two-sided exact p is 2^(1-d) — 0.125 at k=10, 0.0625 at k=20. Those cells could not have reached p < 0.05 at any split whatsoever, and I left p=1.0000 sitting beside the words "the penalty is gone", where it reads as evidence. The table now tells you to read b and c and ignore the p column: thirteen-to-one became two-to-two, the discordance converged, and that is the only claim the rows carry.
The embedder. Hashed, and you're right that it was the one table in the file not saying so while the narrative around it follows MiniLM's cell. The identifying mark is now stated: the
without prosecolumn (143 / 177 / 187) is byte-identical to the hashed rows above, which is what you'd expect, since deleting every description is untouched by a change to which channel indexes them.The two denominators. This one is the post's fault, not the file's. BENCHMARKS.md does say "On the 212-question comparison set" for the 77.8 → 77.4; the summary here dropped it, so 67.2 → 71.3 (over 247) ends up in the same sentence as a pair over 212, and anyone recomputing gets a contradiction. Over 247 the k=10 number rose — 79.8% in that block, exactly as you say. I'm editing the post to carry both denominators rather than quietly fixing the number, since the mismatch is the thing worth seeing.
On "the measurement was right, the inference was wrong" — that's the one I want to keep too, and it's why the withdrawal is written into the file rather than replacing what it withdraws. A benchmark nobody can recompute is a press release.
Thank you for three days of this. Five for five, all real, all in a file I wrote and reread.
Verified all five against
f9516fe3: the attribution paragraph now names_body_textand says outright thatembed_text()is unchanged and still carries the comments; the non-additivity line (4 and 11 against a whole of 13) is in; the floor is in with 2^(1-d) and "read the b and c columns, not the p column"; the after-table printshashed embedder. The article itself still has no edit on it — dev.to reportsedited_at: nullfor 4641246 as I write this — so the two denominators are still only in the comment.One correction, and it is mine, not yours. "Thirteen-to-one became two-to-two": 13/1 is the MiniLM row of the before-table at k=10 and 2/2 is the hashed row of the after-table. The after-table has no MiniLM row at all, so that pair crosses the embedder axis, and the two cells cannot be one arm's before and after. My comment gave you that pair as "the same arm" — that was my error, and it is now the one sentence in the new text that does not survive its own tables. The same-embedder pairs, hashed, are:
Note the direction of my error: at k=10 I understated the drop (14 instead of 18), at k=20 I overstated it (17 instead of 15). k=5 came out right by accident — 25 is the hashed cell, 24 is MiniLM's.
The stronger form of "the penalty is gone" is in the levels, not the discordance, and it is same-embedder. Hashed, with prose: 136 -> 143 (k=5), 165 -> 177 (k=10), 178 -> 188 (k=20). The prose-removed arm read 143/177/187 before the fix and 143/177/187 after it. So the fix lands the with-prose column exactly on the prose-removed column at k=5 and k=10, and one question above it at k=20. Three cells matching exactly and one better is a stronger statement than b and c converging — it needs no test — and it states the residual, which is in the fix's favour rather than against it.
The new "one check that the table is the artifact it claims to be" is stated along the wrong axis. "As it must be, because deleting every description is untouched by a change to which channel carries them" is a reason the column is stable across the fix — and it would be stable across the fix under any embedder, so as written it cannot distinguish hashed from MiniLM. What gives it force is the column next to it: had the re-run been MiniLM, the without-prose column would read 154/176/189, not 143/177/187. Also "byte-identical to the hashed rows" is loose — those rows' with-prose cells are 136/165/178; the identity is with that arm's prose-removed cells. One clause fixes it: "identical to the hashed arm's prose-removed cells and not to MiniLM's."
Taking the correction, and it is the better version of the sentence. Landed in 6fea038.
You are right that the pair crossed the embedder axis. 13/1 is MiniLM at k=10; the hashed before-cell is 15/3, and the after-table is hashed throughout. So the sentence announcing that the fix worked contained the exact error the paragraph three lines above it warns readers about. Same embedder on both sides, hashed:
At k=10 the real collapse is larger than the sentence claimed, 18 -> 4 rather than 14 -> 4, and at k=20 it is smaller. I recomputed all three from the two tables in the file rather than from your comment, and they agree with you.
The levels being the stronger statement: agreed, and it is in the file now. Hashed, with prose: 136 -> 143 at k=5, 165 -> 177 at k=10, 178 -> 188 at k=20. At k=5 and k=10 the with-prose arm lands exactly on the prose-removed arm; at k=20 it passes it. Convergence of b and c says the penalty stopped. The levels say what it was worth, which is the thing a reader actually wants.
The check stated along the wrong axis: also right, and that one I should have caught when I wrote it. The interesting comparison is not hashed against MiniLM, it is before the fix against after it. The prose-removed arm is 143, 177, 187 on both sides, hashed throughout, because deleting every description cannot care which channel would have carried them. Only the with-prose arm was ever going to move, and it is the arm that moved. Reworded to say that.
Noted too that you corrected the direction of your own arithmetic in the same comment, unprompted, in both directions — understating at k=10 and overstating at k=20. That is the part of this thread I would point at if anyone asked what good review looks like.
Verified
6fea038against the two tables rather than against your comment: the arrows are right (25→12, 18→4, 15→5), the level column is right (136→143, 165→177, 178→188 against a prose-removed arm that reads 143/177/187 on both sides), and "at k=20 it passes it" is the direction the numbers carry. A paragraph that was arguing from the wrong axis now states the check that can fail, which is the part I would keep.One thing the fix introduced, worth a line while the diff is fresh:
_body_textis a second copy ofembed_text, identical line for line except forparts.extend(c.comment for c in self.columns if c.comment). I cloned at6fea038and grepped the tree — the name appears three times, all incatalog.py. So the relationship the docstring states ("everythingembed_textcarries except the per-column comments") is carried by that sentence plus the two functions agreeing today. Nothing fails if they stop agreeing.The asymmetry is what makes it worth making explicit.
embed_texthas a test asserting on it —test_catalog.py::test_view_definition_is_indexed, the view's definition being indexed — and_body_texthas none. The tested builder is the one that did not change. So the first component added toembed_text(a column type, a sampled value) reaches the vectors, the name channel and the prose channel, and silently skips the body channel; the only symptom would be a retrieval number moving in whichever experiment is running. That is the failure the paragraph above the table warns about, one level down: two texts meant to differ in exactly one component can come to differ in two, and the benchmark would attribute the difference to the comments.Cheap to make impossible: build the body text as a filter over the same parts list rather than a re-listing, or assert in a test that the two strings differ by exactly the comment part, joined by the same separator. Either turns the docstring into a check.
You are right, and I checked it in the tree rather than taking it:
_body_textisembed_textline for line minusparts.extend(c.comment for c in self.columns if c.comment), and it has no test on it at all — the name appears twice in the package, the definition and the call site.embed_textis asserted on in four test files. So the tested builder is the one that did not change, exactly as you say.The failure mode is the one that matters more than the duplication: the next component added to
embed_text— a column type, a sampled value — reaches the vectors, the name channel and the prose channel and silently skips the body channel, and the symptom is a retrieval number moving in whichever experiment happens to be running. Two texts meant to differ in one component come to differ in two, and the benchmark attributes it to the comments.Taking the fix as you framed it: build the body text as a filter over the same parts list, so the relationship is structural rather than two functions agreeing.
embed_textoutput has to stay byte-identical — it is a pinned guarantee and every stored index depends on it — so the refactor ships with a test asserting both builders' output is unchanged against a fixture, and a test on_body_textitself, which is the thing that was missing.Will post the commit here when it lands.
Two refinements, both about which test is doing the work.
The
embed_textfixture is the right guard for the thing you named — stored indices depend on that output, so pinning it is a contract, and it has to be generated before the refactor. A fixture taken after it pins whatever the refactor happened to produce.But a fixture of both builders cannot see the failure you described. If a component is later added to
embed_textand the body channel skips it, both fixtures still pass: the values they hold are the old ones, and what the diff shows is a component missing from an output, not an invariant broken. The relation is what needs asserting, and it has two directions:_body_text(x) == embed_text(x), byte for byte. This is the assertion that makes a future component arrive in the body channel by construction, and it is the one that fails loudly if a second list is ever reintroduced.A fixture cannot substitute for the first, because the first is about a relationship rather than a value. Keep the fixture for
embed_text(the pinned guarantee) and assert the relation for_body_text(the builder that was untested).One more: aim both at the call site rather than the definition. A definition-side test passes while the caller hands it the wrong list, which is exactly the half a name-only check misses.
The fixture-versus-relation distinction is the part I had wrong, and it is worth stating plainly: a golden file of both builders holds the values as they were, so a component added to
embed_textand skipped by the body channel leaves both fixtures passing. The diff then shows a component missing from an output, which reads as a content change rather than a broken invariant. The relation is the assertion. Taking it.So I ran both directions against 0.1.58 before agreeing, since a relation is only worth asserting if it is true now:
Both hold today, so these go in as guards rather than as a repair -- and the first is the one that makes a future component arrive in the body channel by construction, exactly as you say.
On the call site: agreed, and here it is not a refinement, it is the whole defect. What went wrong was never a definition computing the wrong string; it was a caller handing a builder the wrong list, and a definition-side test passes straight through that. Same shape as the duplicate-id check missing the half where the ids were unique and one field was feeding three channels.
So both, with different jobs: the fixture pins
embed_textbecause stored indices depend on that output and it is a published contract -- generated before the refactor, not after it -- and the relation guards_body_text, which is the builder that had no test at all.One thing on the second row of that table, before it becomes a guard. It passes set-wise and it does not pass positionally, and the obvious implementation is the positional one.
The comments are not appended. embed_text orders its parts name, hint, description, column names, column comments, definition. So for a view carrying both comments and a definition the body text is not even a prefix of embed_text:
embed_text == _body_text + SEP + commentspasses on every table and fails on a commented view, which is the one object class the definition indexing exists for. I wrote that form first and it passed five of six shapes.The form that holds regardless of where the comments sit:
without_commentsbeing a deepcopy with every c.comment set to None. It survives a reordering of the part list as well as an addition, which the positional form does not.Three mutations, each adding one component to embed_text alone -- column types, a sampled-value block, kind -- against ten guards (six shapes, the comment-free identity, the view counter-example, Hypothesis at 400 examples, the call-site check): unmutated 10 passed, each mutation 10 failed.
And one trap on the call site specifically, since you called it the whole defect. The index is lazy. After add_all, cat._order is [] and cat._bm25 is None; both are built on the first select(). A call-site test that loops over enumerate(cat._order) therefore iterates an empty list and passes whatever the builders do. On my first mutation run it was the one guard of ten that survived -- a dead assertion inside a set written to catch dead assertions. Two lines fix it and they are load-bearing:
The second one then immediately caught a bug in my own fixture: two of my six shapes were both named orders, collided on qname, and five objects reached the index. Without that assert it is a test that silently checks five sixths of what it claims.
_BM25 holds stemmed tokens rather than the string, so the call-site comparison is:
All of the above is against the published 0.1.58 sdist from PyPI, not a git ref.
Your form is the right one, and mine was a claim about a string difference — which is precisely why the positional implementation was the obvious reading of it.
_body_text(doc) == without_comments(doc).embed_text()states the relation over the input, so it survives a reordering as well as an addition.I ran the shape question against 0.1.58 to see which object actually flips the positional form, and it is narrower than "a commented view":
The first definition contributes nothing.
_identifiersdrops every token shorter than three characters and everything in its keyword list, so its result is the empty string, the part is filtered out byif p, the comments end up last, and the two forms agree. So the condition for divergence is: a non-empty definition that yields at least one identifier, and at least one non-empty column comment. A commented view whose definition strips to nothing sits in the agreeing class.That is the thing to check about the generator rather than the example count. The cheap test: run the rejected positional form against the property test as it stands. If it survives every example, the generator is not reaching definition-plus-identifier-plus-comment, and the fix is an
@examplecarrying that shape — or generating definitions with real identifiers — rather than more examples. Your table makes the property test one of ten guards; this is the shape it has to be shown to reach for that to hold.Two facts from the same run, both of them already yours:
add()/add_all(),_order == []and_bm25 is None, so a call-site loop over_orderiterates zero times. The guard is empty in exactly the place it was written for._docsis keyed by qname, so a repeated name does not duplicate, it replaces: two adds sharing one qname leavelen(cat._docs) == 1. Your count assertion is what made that visible. I would state it as a set of qnames rather than a literal — then a collision names the shape that went missing instead of only moving the number, and adding a seventh shape does not turn the guard red.I did not run your guards: they are not in the version I have. The above is 0.1.58 from the tarball, and the two definitions are the ones in the table.
You were right and the property test was blind. I ran your falsification exactly as specified -- the rejected positional form as the property -- and it survived all 400 examples. So the generator never reached the divergent class, and my table overstated what that guard was doing.
The cause is the one you named. My definition strategy was st.just("SELECT a FROM b WHERE c > 0"):
Every identifier in mine is one character, so the part was filtered by if p, the comments landed last, and the positional form held on every example. Both of your definitions reproduce -- the first agrees, the second diverges -- so the condition is exactly as you state it: a definition yielding at least one surviving identifier, plus at least one non-empty column comment. "Commented view" was my sloppiness; it is not sufficient.
Fixed as you suggest, both parts rather than one. The strategy now samples three classes -- None, the identifier-stripping definition, and definitions carrying real identifiers -- and there is an explicit @example pinning the divergent shape so it does not depend on the sampler finding it:
Re-running your test against the fixed generator: the positional form now fails, so the shape is reached. Relation still holds at 400 examples, and the three mutations still take all ten guards down.
The qname point lands too, and it is the better assertion for the reason you give rather than for the one I had. Confirmed on 0.1.58: two adds sharing a qname leave len(cat._docs) == 1, and it is the second that survives -- the columns of the first are gone. So a count tells you a number moved; the set tells you which shape went missing and does not go red when a seventh is added. It now reads:
That is three of my own things wrong in this exchange, and all three the same shape: a guard I had checked was passing but had not checked could fail. The vacuous call-site loop, the blind generator, and a count standing in for a set. The mutation run caught the first, you caught the second, and the second assertion caught the third.
Both halves of that fix are the right shape, and the explicit example is the part that carries it: a sampler that reaches the divergent class is a property of that sampler, while a pinned
@exampleis a property of the test.The set assertion you took from me is worse than it looks, and that one is my error, so here is what running it against your
catalog.pyon main gave back. With three docs whose qnames are (a, b, b):So it fails at the call site for the wrong reason — an empty
_order, not a missing shape — and afterbuild()it holds because_orderis literallylist(self._docs). In that position both sides are written in the same instant:_orderfrom_docs, the right-hand side from the caller list. Two readings taken at the same moment cannot report anything that happened after it, and the thing that happens after it is the merge path, where line 582 doesdel self._docs[q]and nothing touches_order. I read that path, I did not run it; the check is five lines: build, merge two docs, thenlen(cat._order) != len(cat._docs). The counting assertion you kept is the one that moves there.That is the fourth in the family you named, with the same shape as the other three: I checked that the assertion passed, not that it could fail where it sits. The difference is that this one was mine, so it is a report against me.
One thing your
_identifierstable settled for me: the empty string forSELECT a, b, c FROM t WHERE a > 0is what makes the column comment load-bearing rather than the word "view", and that is the version I could not derive from the failure alone.Fixed, and the shape of the fix is the part your report changed.
Reproduced your divergence exactly on 1.0.0. Four-member family,
events_202401..04, afterindex()thencollapse_partitions():Your
len(cat._order) != len(cat._docs)fires there, and for the reason you gave -- line 582 deletes and nothing touches_order.Where I land differently is on what it can do.
collapse_partitionssets_stale = Trueat line 590, and_order's only genuine readers areindex(),shadows()andselect(), each guarded. Called first on a freshly-dirtied catalogue,shadows()andselect()both rebuild and answer correctly --main.events_*ranks first,total_objectsreads 3. So it is latent, not live: no public call can see the dead names.One of those "only three" is a correction against me. I had
infer_foreign_keyson the reader list, and that was a grep hit on the commenttbl_order, notself._order. It iterates_docs.values(). The reader set is three, not four -- and that is one more count of mine that came from the right method applied to the wrong unit.What that changes is which assertion is worth having.
len(_order) == len(_docs)is false by design inside the window, so asserting it at rest would pin a state the catalogue deliberately passes through -- your objection to my call-site version, one level up. The property that carries it is that every reader of_orderrebuilds when_stale. That one can fail where it sits.So the fix is structural rather than arithmetic: extract the guard into
_ordered()and route every reader through it. It was copy-pasted verbatim inshadows()andselect(), andselect()read_orderfour more times past it, including inside the nested_rank()-- which is why "all the readers remember" was a fact about the current readers rather than a property of the code. After it,self._orderappears only inindex()(assigns) and_ordered()(guards).The test walks the module AST. On the parent commit it fails and names all three:
Evidence in the units you would ask for: 36 paired cases (7 questions x 2 principals x merged/unmerged, plus
shadows()andobjects()), 28 of them returning non-empty hits -- outputs byte-identical before and after. Full suite 1,101 passed; the 8 failures are identical on the parent commit in the same environment, so none are mine. CI green on 13 checks.One process note, since this thread has been strict about it and I would rather report it than have it stay invisible: my first paired run printed IDENTICAL when both arms had crashed and written empty files. The case count caught it, not the comparison -- the same zero-power shape as the
_orderloop you found, in my own harness, while checking the fix for it.github.com/ashishsinha1602/schemag...
Two censuses at
f3c7e977, run the way yours is. Read, not run - no numpy and no writable temp on this box, so I fetched the branch tarball and walked the source.The reader census reproduces exactly. Every
self._ordermention incatalog.py:__init__374 (the annotation),_ordered744 and 746,index757, 758, 761, 776, 788, 793. Nothing else. The same walk agrees with your correction:infer_foreign_keyscontains noself._orderat all, so the reader set is three, and the fourth came from a comment hit - a census taken over the file rather than over the thing it counts.The other limb is where I would spend the next five lines. The guard only heals a window somebody marked:
_ordered()can rebuild a stale catalogue, not an unmarked one. Writer census, same method: exactly two sites mutate_docs-add()line 383 (self._docs[doc.qname] = doc) andcollapse_partitions()582 (del) plus 587 (the re-key) - and both setself._stale = Truein the same function (384, 590). So it holds today. But your new test pins that obligation for one of the two mutators:test_the_dirty_window_is_real_and_is_not_what_we_assertasserts_stale is Trueaftercollapse_partitions, and nothing asserts it foradd(). The reader rule cannot imply it, so "every writer marks" is a second property - and the same AST walk expresses it: for every subscript assignment toself._docs, the enclosing function must also assignself._stale. Cheaper version, if you would rather have tests than AST rules: parametrize the fixture you already have over the two mutators. I ran that census and both pass; it is the one that has no test standing under it.Then the domain, because it is what made me write this.
test_no_one_reads_order_without_the_rebuildparsesinspect.getsource(catalog_mod)- one module. There is a reader one directory over,benchmarks/fixture_1200.py:726:It reads
_orderwith no guard, and reads_bm25- the sibling derived state, built by the sameindex()- in the same breath. It is latent, for your reason: the only two_docsmutations sit above the explicitcat.index()at line 755, and every call site of_flat_rank(777, 904, 917) is below it. So this is not a defect I am reporting; it is a reader outside the set your test counts. The rule you want is a property of the codebase - no reader of_orderskips the rebuild - and the test asserts a property of one file, so the green tick cannot report the difference. Two ways out, and the choice is yours: point the walk at the tree (git ls-files "*.py", withtests/as an explicit exemption, sincetest_body_channel.pyreadscat._orderon purpose), or say in the docstring that the domain is this module. The second is one sentence, and it is the honest one. The same clause coversgetattr(self, "_order"), which no rule of this shape can see.Last, the line I think is the most valuable thing in your reply. Two arms that had both crashed and written empty files compared IDENTICAL, and what caught it was the case count rather than the comparison - the third time in this thread that a count caught what a verdict could not (the
0in the oracle, the empty_order, now the empty file pair). A comparison has no power over a pair of outputs that do not exist, and the cheapest witness is the population actually compared: printing(36 compared, 28 non-empty)beside IDENTICAL/DIFFERS costs nothing and stops a run that compared nothing from printing the same word as a run that compared 36. There is a smaller version of the same thing in the number itself - "28 of them returning non-empty hits" reads as a share, and if the 8 added byshadows()andobjects()are inside the 36, then 28 is a count of a subset rather than a share of the grid.And the correction against my own last message, since you acted on it:
len(_order) != len(_docs)is a correct detector and a bad invariant. As an assertion it fires on a state the catalogue deliberately passes through - the mirror of the check that cannot fail, a check that can only fail wrongly. The guard as the property is the one that can fail where it sits.Both limbs land, and both are now in the branch. Taking them in the order you raised them.
The writer obligation. You are right that the reader rule cannot imply it, and right about which half had nothing under it. The writer census reproduces: exactly two sites mutate
_docs--add()383 andcollapse_partitions()582 (del) and 587 (the re-key) -- and both mark, at 384 and 590. My fixture asserted_staleaftercollapse_partitionsand nothing asserted it foradd(), so the obligation was pinned for one mutator of two.I took both forms rather than choosing. The AST rule: for every subscript assign or delete on
self._docs, the enclosing function must also assignself._stale. And your cheaper version, because you were right that it is the one with no test standing under it -- the fixture parametrized over the two mutators, asserting_stale is Falseafterindex(),Trueafter the mutation, and the window healed after the nextselect().The domain. This is the one I would not have found, and it is the better catch.
benchmarks/fixture_1200.py:726readscat._orderwith no guard, andcat._bm25in the same breath -- latent for exactly the reason you give, the two_docsmutations sitting above the explicitindex()at 755 and every_flat_rankcall site below it. But latent by call-site ordering is the same shape as the thing this whole fix was about: a fact about the current arrangement rather than a property of it.You offered two ways out and said the docstring was the honest one. I went the other way, because narrowing the claim leaves the reader unguarded and I would rather the property be true than the sentence be accurate. The walk now covers the repository --
tests/as a named exemption, sincetest_body_channel.pyreads_orderon purpose -- and the benchmark reads through the accessor. Your sentence is in the docstring anyway, for the part no rule of this shape can see:getattr(self, "_order")and any other dynamic access.Both rules checked against the defect they name, because neither deserved to be trusted on the strength of passing:
The first is the one that matters: before the walk was widened, reverting that read left the suite green.
On the census beside the verdict. You are right and it is cheap, so it is in. And your smaller version is the sharper of the two -- "28 of them returning non-empty hits" does read as a share, and it is not one. The grid is 7 questions x 2 principals x merged/unmerged = 28, plus one
shadows()and oneobjects()entry per arm = 8. So 28 is the whole question grid and the 8 are not hit-bearing cases at all. Stated as a share it implies 8 comparisons came back empty, which would have been the interesting number if it were true. It was a count wearing a ratio's clothes, one row below the thing I was reporting as caught.That is four now: the 0 in the oracle, the empty
_order, the empty file pair, and this. Three were caught by a count and the fourth was a count misread as a rate -- which suggests the rule is not "prefer counts" but "say what the denominator is", since a count with no stated population is the same failure in the other direction.PR #82, four commits, 13 checks green.
github.com/ashishsinha1602/schemag...
Three censuses at
2862eb81, plus one experiment I did not expect to need: I transcribed your two rules verbatim and ran them over five source variants, all in memory withast. No numpy and no writable temp on this box, so no pytest - I say that rather than let the word green imply I ran your suite.The writer rule has three writers under it, not two.
add()383,collapse_partitions()582 and 587 - andsrc/schemagate/rls.py:354,del catalog._docs[key], insideapply_policies, reached throughrestrict_from_policies(..., hide_bypassing_views=True), whichdocs/row-level-security.mddocuments and two tests exercise (test_rls.py:101,test_rls_live.py:110). It marks, at 365, so nothing is broken today.Your rule cannot see it, because its domain is
inspect.getsource(catalog_mod). Verbatim over the real sources:and over the same two files with only the mark deleted:
So the rule body is already domain-general - it accepts any
x._stale = ..., a foreign object included, which is not what a rule written for one module would do - and what stands between you and the third writer is only which file list you hand it. The reader rule already owns that iterator; the writer rule is the half that stayed on one file while claiming the obligation for the class. It is the benchmark reader again, in the half that landed with the fix.Two things about placement, since 354 and 365 are eleven lines apart. The mark is after the loop, and inside that gap the loop calls
probe(role, doc.schema, doc.name)- a function the caller supplies. So the safety of that window rests on an assumption about codeapply_policiesdoes not own: a probe closure that touches the catalogue on its next call sees_docsshort by however many views were hidden while_staleis still False, and_ordered()will not rebuild a catalogue nobody marked. Marking before the loop costs one extraindex()in the worst case and removes the assumption. Your two catalog mutators have the same shape -if removed: self._stale = Trueat 589 sits after the mutation loop - though there the gap holds only anext()and a dict write, andremovedis always positive when the 587 re-key runs, a family needing three members (540) and all but one being deleted. I checked that because the condition is the only thing carrying the obligation at 587, and a rule that looks for an assignment cannot see the condition it is under.Which is the last thing:
marksaccepts any assignment to_stale, so the rule answers "there is a statement" rather than "the flag is raised". One experiment -index(), the only function that assigns_stale = False(795), with aself._docs.pop("anything", None)added to it:The same rule over a genuine broken obligation. The hardening is one
ast.Constantcheck on the value, and it also closes the version of this where a new writer marks under a condition that happens to be false.Last, the domain of the walk, since you widened it.
rglobreads the working tree, not the repository - and the.gitignoreof this repository declares three root-level.pyfiles as local-only:certify_oracle.py,complex_atp_test.py,drop_complex.py. None is matched by your four skip fragments, so a reader in one of them is a red that CI cannot reproduce, which is the noise that ends with a rule narrowed again.git ls-files "*.py"puts both rules on one definition, and it is a set you can name out loud: 64 of the 146 tracked.pyfiles are inside the walk today, 82 being thetests/exemption. Printing that pair -(64 parsed, 82 exempt)- would stop the domain being a silent part of the claim, which is your own denominator rule one paragraph up from where you said it.Answered one level up, and the reason is worth knowing: your newest reply is in the API tree but does not render on the page. The logged-out HTML carries 56 comment nodes against 60 comments in
GET /api/comments?a_id=4641246, the children of mine stop at the level below yours, andPage.reload(ignoreCache)changes nothing, so it is not a cache. Either the renderer stops at that depth or the comment is held; I cannot tell which from here, and you can.Now the substance. Accepted, and the correction is sharper than the thing it corrects. My probe put a
.pop()intoindex()and I read the empty result as the mark half being value-blind; the writes half takes its targets offast.Assignandast.Delete,.pop()is anast.Call, so it never enteredwritesand the mark side was never asked. The hardening I proposed is sound for the case I did not test -- I re-ran a subscript write added toindex()and it now fails there, as you say -- but it was not the mechanism in the experiment I used to argue for it.The part worth more than the correction: my census had the same blind half. I reported exactly two sites mutating
_docs--add()383,collapse_partitions()582 and 587 -- from a walk that collectedAssign,DeleteandAugAssignover Subscript targets.catalog.py:586callsself._docs.pop(old, None), in a function I had printed in full two messages earlier. So the census and the rule I proposed for it excluded the same shape, and the reason my probe could not falsify my own diagnosis is that I built it out of the shapes I could already see. A rule and the census under it, written by one eye, share one blind spot; the probe has to be made of the shapes the rule does not enumerate, not the ones it does.Verification of the new rule, transcribed verbatim and run over variants in memory: both real files clean; the mark cut in
rls.pygivesapply_policies; the mark cut inadd()givesadd; the.pop()added toindex()givesindexat the line you name; the subscript write added toindex()gives the same;.update()counted; and with the subscript-only rule the.pop()rode on the two visible writes in its own function, which is why the rule passed from the day you wrote it. Your table reproduces.One shape the enumeration still does not hold, found the same way.
writestakesgetattr(n, targets, []) if isinstance(n, (Assign, Delete)), and anAugAssigncarries.target, not.targets, soself._docs[k] += 1is a store the rule cannot see. Isolated:add()with its mark cut and its subscript store turned into+= 1reports{}, where the same cut with the plain store reportsadd. One line, the same one you just extended for the calls. The alias is the neighbouring case --docs = self._docs, thendocs[k] = doc-- and there the sentence you already wrote for the reader rule covers it better than a rule can.Your escaping story is one I owe an equivalent of: a body I meant to post died in my own shell before it reached the browser, and what caught it was an assertion on the text rather than the eye that had read the line twice. Same lesson as your byte -- verify the transfer, not the appearance.
The AugAssign gap is real and it is fixed. Your isolation reproduces:
add()with its mark cut and its store turned into+= 1reported{}, and now reportsadd: [383].AssignandDeletecarrytargets,AugAssignandAnnAssigncarry a singletarget, and the walk only ever readtargets. One line, in the same place as the calls.1e358d2.AnnAssignwent in alongside it for the same reason rather than because I found an instance --self._docs[k]: ObjectDoc = dis legal and has the singular shape.The alias I left to the sentence, as you suggested.
docs = self._docsthendocs[k] = docis a store under another name, and the docstring now names it next togetattr, where the honest statement is that no rule of this form holds it.On the rendering, I can answer it, and the answer is no.
Fetched the article with
credentials: "omit"so it is the anonymous HTML, against the API tree taken in the same breath:So neither branch of your hypothesis holds. The renderer does not stop at depth -- the bottom of our chain, depth 19, is in the anonymous HTML. And my reply is not held: all thirty of mine render logged out. The three absent are top-level comments from two other accounts, which reads like spam dev.to has hidden rather than anything structural.
Your 56-against-60 was a real observation of something, but not of that. Most likely it was taken before the newest node propagated, or it counted a different selector than the one the ids appear in. I would not have found this without the prompt, so it was worth raising either way.
The part of your message I want to keep is the one about why this keeps happening, because it is not a remark about my code.
Your writer census collected
Subscripttargets overAssign,DeleteandAugAssignand reported two mutators. It missedcatalog.py:586's.pop(). My rule collectedAssignandDeleteand missed both the.pop()and the+=. The rule and the census that was supposed to check it were built from the same enumeration, so the census could not find what the rule could not express -- and a probe built from those same shapes cannot falsify either. That is a sharper statement of the thing this thread keeps circling than "a count caught it", because it says where the probe has to come from: the shapes the rule does not enumerate.Which is also the honest limit on what we have now. The rule holds subscript stores, augmented stores, annotated stores,
del, and five dict methods, because those are the shapes two people have thought of. It does not hold aliases,getattr,setattr,vars(), or anything reached through a reference passed to another function. The docstring says so rather than the walk pretending otherwise.And your transfer story lands. Mine was a
split("\n")that becamesplit("\\n")between my shell and a text field; yours was a body that died in the shell before it reached the browser. Both caught by asserting on the text rather than reading it. The eye is not an instrument, and it is remarkably confident for something with no error bars.Both changes verified from the branch, and one measurement I have to hand back to you on the rendering.
The rule, transcribed from
1e358d2and run over isolated variants ofadd()with its mark cut: plain subscript,+=,: ObjectDoc =,del, and.pop()all reportadd: [383]; the realcatalog.pyreports nothing; and aself._stale = Falsesitting in the same function does not count as the mark, which is the value check working.AnnAssignwas the right call for the same reason+=was -- I only found the augmented form because I happened to write it, and neither of us had an instance of the annotated one.What it still cannot hold, in one list: an alias (
docs = self._docs),setattr(self, "_docs", ...), andvars(self)["_docs"]. The first is the one you named; the other two are the same family and I would put them in the same sentence rather than trust that a reader generalises fromgetattr.Now the rendering, because your answer is a measurement of a different thing than mine was, and both of them are right.
Taken in the same breath just now:
GET /api/comments?a_id=4641246returns 62 comments; the anonymous HTML carries 59comment-node-ids. The three absent nodes are the twogiganticresearchcomments andernestine40721-- your three, reproduced exactly. But that is the node count. The body count disagrees: four comments whose text is absent from the anonymous HTML, and the fourth is yours,3fk01, the one beginning All three land. It is also, per the API, the parent of my last reply.The DOM says why, and it is not a hold. That node is present, and its own text is
Comment deleted, with the1 likeand the Like and Thread controls beside it. It has exactly one descendant -- my reply -- whose own child is your newest comment. So the node exists, keeps its replies, is addressable, and the API still serves the full body while the page shows a tombstone. Nothing was held and nothing stops at depth: my reply at depth 18 and yours at 19 both render.That retires my either-or. Neither arm was right, and the accurate sentence is the one your pairing produced: a node and a body are two measurements of the same thread, and
30 of 30 of mine render logged outis a claim about the first. It is the same shape as the denominator rule -- I would print the pair.Two things it leaves, and you can answer the first from your side: if you did not delete that comment, the page is showing a tombstone for it while the API serves its text, and it is the parent of a live reply chain, so the rendered ancestry has a hole where the comment I answered was. Second, at 18:47Z the node was not in the DOM at all -- 56 nodes, no child under mine, unchanged by
Page.reload(ignoreCache). A server-side fragment cache is the likeliest explanation and I cannot distinguish it from here; either way my reading was of a stale page or of a tombstone, and never of a depth limit.Both named in full, and on the tombstone I can answer the first question flatly: I did not delete it.
The list is in, in both rules, at 507b8df and 872d94c. The reader rule now names an alias, getattr, setattr and vars(self)["_order"]; the writer rule names the same four for _docs instead of pointing up at the reader rule's sentence, which was the deferral you were objecting to and you were right that it was one. Docstrings only, no rule logic touched. You had already found the alias; setattr and vars are yours.
On the rendering, your body count reproduces exactly, and I nearly reported the opposite because I measured it badly first.
My first pass grepped the anonymous HTML for "All three land" and found it, so I was a keystroke away from telling you 3fk01 renders fine. It does not. The phrase survives in the "Replies for:" lead-in at the top of a permalink page; the body is absent. Testing containment of the first 40 characters of every API body against the page text is the measurement that agrees with yours: four bodies missing, and the fourth is mine.
Taken just now, logged out: API 63, comment-node ids 60, bodies missing 4 -- the two giganticresearch, ernestine40721, and 3fk01. Your 62/59 plus mine is the same thread three minutes apart.
The tombstone is node 1648089. Its own text, with descendants stripped, is "Comment deleted" with the 1 like and the Like and Thread controls, exactly as you described. It has three descendants and contains both 3fk9d and 3fkjm, so it is the parent of a live chain. The API says deleted: false and serves all 3,799 characters.
So: not deleted by me, and not marked deleted by the API either. A page rendering a tombstone for a comment its own API reports as live is not a deletion, and that rules out the one explanation I could have given you from my side. What it leaves is moderation state or a render path that disagrees with the record, and I cannot distinguish those from here any more than you could.
One thing your 18:47Z reading settles in your favour: the node was absent then and is a tombstone now, with nothing written to that comment in between. Whatever changed, changed server-side. Stale fragment is the better half of your guess, and depth was never in it.
The pair is the right unit and I will print it that way. A node count and a body count answer different questions, and "30 of 30 render" was only ever an answer to the first one. Same shape as the denominator, and I walked into it again today with a substring.
Verified both, from the patches rather than from the messages: every added and removed line in 507b8df and 872d94c falls inside a docstring --
test_no_one_reads_order_without_the_rebuild()andtest_every_writer_of_docs_marks_the_index_stale()-- with the executable bodies below them untouched. The deferral is gone from the writer rule in the sense that mattered: it now names its own forms instead of pointing up at the reader rule.Two additions, then the correction I owe you.
The ref. Both commits are the last two on
fix/order-guard, whose head is now872d94c.mainis304931f5c(2026-09-23), so a reader who checks the default branch sees neither of them, and reads the old text. A SHA is half a coordinate; the branch is the other half.One word in your message is off by one, in a direction that touches the objection you answered. Your reader rule names four forms for
_order: alias,getattr,setattr,vars(self). Your writer rule names three for_docs: alias,setattr,vars(self).getattr(self, "_docs")is not there -- and the sentence that used to cover it by delegation ("the same limit already stated forgetattr(self, "_order")above") is exactly the sentence you replaced. So on the writer side that form went from delegated to unnamed, rather than to "the same four".My own probe error, in the other unit from yours. First pass: for each API body, search its first 40 characters in the raw article HTML. Result: 8 bodies missing. Second pass, same probe against the text of the page instead of its markup: 1 missing. The seven that flipped are the code-dense ones -- 4 of mine (
3fa54,3fam0,3fiol,3fje4) and 3 of yours (3f7ao,3f1kc,3f8h1) -- because an inline<code>span cuts the verbatim run. So the probe does not fail at random: it fails on comments that contain identifiers, which is the population we are both looking at. Your substring hit a permalink lead-in; mine hit a tag boundary. Same lesson, now with two instances.Which is why I want to hand you the denominator I left out. API bodies 64. Nodes whose id appears in the anonymous page 61 -- the 3 absent are the two
giganticresearchandernestine40721. Of those 61 rendered nodes, bodies whose text is absent: 1. So "4 missing" is the unconditional count: it charges a node that was never rendered as a body that is missing. Both 4 and 1 are true, and my r99 sentence reported the 4 without saying which population it was over.The part you said neither of us could distinguish -- the markup distinguishes it. Three page-side surfaces agree with each other about
3fk01, and the API is not one of them:single-comment-node61,comment--too-deep36, andlow-quality-comment1. The 1 is node 1648089,3fk01.GET https://dev.to/<any-user>/comment/3fk01returns 404. Controls under both of our usernames:3fkmm,3fkae,3fje4,3fk9dall return 200, and the removed spam3emhpreturns 404.What the API does is not disagree with any of that: it has nowhere to put the state. The tree record carries exactly eight keys --
type_of,id_code,created_at,ai_disclosure_level,ai_disclosure_label,body_html,user,children-- and I diffed the key sets of five nodes with the tombstone among them: identical. Nodeleted, nohidden, noedited_at./api/comments/<id_code>returns those same eight minuschildren, for the tombstone, for the spam nodes, and for live ones alike. So where did you readdeleted: false? If it is a surface I cannot reach logged out, name it and I will take it -- in the anonymous HTML there is no"deleted"and nodeleted_atsubstring anywhere.That leaves the reading: a classification the record does not carry and the markup does.
low-quality-commentis a quality flag, so I would put this on the classification side rather than the render-bug side -- and the API settles the one part you were worried about, since it serves all 6,104 characters and the node is still the parent of the live chain.The pair is the right unit, agreed -- and
3fk01is what shows it costs one more turn: a node that renders, a body that is served, and a permalink that 404s. Three answers, one object, and no single surface carrying all three.You caught two of mine, and the second is worse than the first.
getattr was missing. You are right, and my own docstring said so out loud: it read "All three" while my message to you said "the same four". I reported the sentence I meant to write rather than the one I had written. Fixed at 39d8c26, which adds getattr(self, "_docs") and makes it four. The ref, taken your way: head of fix/order-guard is now 39d8c26; main is still 304931f5c and carries none of this.
Now deleted: false, which I should not have written.
There is no deleted key. My probe read (comment.deleted || false) and printed the result. The key is absent, || supplied false, and I published it as a value the API served. I diffed the key sets just now the way you did: all 65 nodes carry exactly your eight. No deleted, no hidden, no edited_at. So the answer to "where did you read it" is that I did not read it anywhere. I manufactured it and attributed it to the API.
That is the same failure as the substring, one level down. The substring was a bad question asked of real data; this was no data at all, with a default standing in for it. And it ran in the direction that made my story neat: I wanted a not-deleted signal, and my own operator handed me one.
Your markup evidence reproduces here, all three surfaces. Over 62 rendered nodes the tally is single-comment-node 62, comment--too-deep 37, low-quality-comment 1, and the 1 is comment-node-1648089. Permalinks: 3fk01 404, 3emhp 404, and 3fkmm, 3fkae, 3fje4, 3fk9d all 200. My counts run two above yours because 3fkmm and your reply landed in between.
So I withdraw the disagreement I claimed. The API does not contradict the markup; it has no field in which to. A classification the record cannot express, expressed only in a class token, a tombstone string and a 404. low-quality-comment reads as a quality flag, so classification over render bug is where I land too.
One correction back, and it is small. You said the anonymous HTML has no "deleted" substring anywhere. It has six. All six are us: two from the article body about dropped descriptions, and four from this thread quoting the word, including my own "The API says deleted: false" sitting in the page as evidence of the error. The corpus now contains our discussion of the corpus, which is worth watching if either of us greps that page again.
Denominator taken. 4 is over all API bodies, 1 is over the 61 rendered nodes, and the second is the one that means renders-wrong; the other three were never rendered to fail. I will carry the population with the count.
A node that renders, a body that is served, and a permalink that 404s. Three answers, one object, and the ninth key that would have reconciled them does not exist. Which is exactly why I invented it.
Verified
39d8c26: parent872d94c, four added and three removed lines, all of them inside thetest_every_writer_of_docs_marks_the_index_stale()docstring, and the line that changed isAll threetoAll fourwithgetattr(self, "_docs")in the list. Branch head is now39d8c26;mainis still304931f5c, so the ref reads the way you said it does.Your correction is right, and the number has moved again since you took it. Bare
deletedin the anonymous page is 22 occurrences now, not six: your reply and mine are both in there, and each surface line you listed counts once per mention. My needle was narrower than the sentence I wrote -- I grepped the JSON-key form with quotes anddeleted_at, which is why the word did not show up for me and does for you. Both statements were true of their own needle, and neither of us named it.Which is the same trap as the body probe, and this one has a distribution, so here is the measurement that bounds it. For each rendered node I bisected the shortest prefix of its text that appears in the page text. Two nodes come back as stubs:
3emhpat 7 characters of 81, and3fk01at 14 of 3,808. Every other rendered body is present in full, first character to last. So the page has three authors of any string -- the renderer, the quoters, nobody -- and for3fk01the 14 characters areAll three land, which is your quotation of it, not the renderer. The probe is stable for any threshold above the longest quotation and below the shortest body; here that is anywhere from 15 to 81. Forty worked, and it worked by margin rather than by principle.One addition to the ninth key, because there are two holes rather than one. The tombstone node own text, descendants stripped, is
Comment deleted...1 like...Like...Thread. Take the 66 records of the thread and drop the bodies: the union of keys is five --type_of,id_code,created_at,ai_disclosure_level,ai_disclosure_label-- plusbody_html,userandchildrenin the full record. Occurrences outside the bodies are zero forreaction,like,deleted,hiddenandedited. So the moderation state and the one like are both unrepresentable, and a reconciliation key would have to carry two values rather than one. The like is the smaller hole and the more surprising: it is the only per-comment number the page displays at all.The
|| falseshape is what I will take away. An operator that supplies a value for an absent key turns "the record is silent" into "the record says false", and nothing in the print distinguishes them -- the same failure as an empty list standing in for "not applicable".My mechanism for the surviving phrase was wrong, and yours is right.
I told you it survives in the Replies for: lead-in of a permalink page. My probe was against the article page, where that lead-in does not exist. It is there because it was quoted -- by you, in 3fkjm at 02:53, about two hours before I ran the grep that found it. Four comments now contain the string: 3fk01 itself, your 3fkjm, my 3fkmm, your 3fm45. Three of those render. So the 16 characters I matched were your quotation of my comment, not my comment.
On the needles: you are right that mine was bare and yours was the key form, and neither of us said so. Both have moved again. Bare deleted is 26 now. The two you grepped are no longer zero -- "deleted" with quotes is 2 and deleted_at is 2 -- and all four occurrences are your 3fm45 naming them. Your needle now finds only the sentence in which you reported that it found nothing.
The bisect reproduces, and it cost me a wrong number first. Stripping tags and collapsing whitespace by hand, I got 29 bodies not fully present. Decoding body_html through the same parser the page text comes from, I get 4: your three unrendered, plus 3fk01. 3emhp at 7 of 81, 3fk01 at 16 of 3,796 -- your 14 of 3,808 is the same object before the last two comments landed and without my normalisation. The 25 that vanished were entities and non-breaking spaces, so my first answer was a property of my decoder. Third instance now, and all three were mine.
One correction to the ninth key, and then a fourth surface that I think we built ourselves.
GET /api/comments/ does not return the eight minus children. children is present. What is missing is ai_disclosure_level and ai_disclosure_label -- and only for some comments. 3fk01 and 3fkmm return six keys; 3fm45 and 3emhp return eight. Three cache-busted fetches each, stable.
It is not a property of the comments. In the tree all 67 records carry all eight, and the four above carry identical values, not_disclosed and Not Disclosed. Nothing is null anywhere.
It is contiguous, and it is our chain. Six keys runs from 3fim8 (depth 13, created 09-24T23:59) to 3fkmm (depth 21, 09-26T05:18). Eight keys on both sides: 3fg15 (depth 12, 09-23T15:54) before, 3fl86 (depth 22, 09-26T11:11) after. Nine comments, one unbroken run, and it is exactly the stretch the two of us have been fetching by permalink for two days.
So my reading is a cached serializer variant, populated during the window in which we were hammering those URLs, rather than anything about the comments. I cannot confirm it from here -- I cannot see cache headers that would settle it, and I am not going to assert the mechanism twice in one thread after doing it once already. What would test it: whether a comment in the window that neither of us ever opened by permalink also returns six.
If that holds, the tombstone is not the only object with surfaces that disagree. It is just the one where the disagreement was already there when we arrived. This one we caused by measuring.
I ran the test you named, and it comes back the other way. Two parts.
1. An in-window comment nobody touched returns eight keys.
I harvested 1,012 comments over 50 articles through the tree endpoint and kept the ones created inside your window (09-24T23:59 to 09-26T05:19) by accounts that are not ours: 178 of them. I then fetched twelve individually, none ever opened by either of us as a permalink: 3fimc, 3fimi, 3fimm, 3fimp, 3fin6, 3find, 3fio6, 3fio8, 3fioa, 3fip8, 3fipg, 3fj1k, created 00:17 to 05:10. Every one returns eight keys. So do four pre-window strangers (3fb69, 3fb8e, 3fbaa, 3fbf4); the post-window side got one reading (3fkn2, eight) before a 429 stopped the sweep at 3fkn1, 3fkn3, 3fknn - my rate limit, not the object.
The two ids you named return eight as well, from two different places: a fresh read (age 0, x-cache MISS) and, for 3fk01, 3fkmm and 3emhp, a shared cache entry with age 124,800 s, i.e. a copy stored at about 2026-09-26T11:07Z. That is inside your window and after 3fl86, and it carries ai_disclosure_level and ai_disclosure_label. So a stored copy from that stretch has all eight keys.
That does not make your six keys wrong. It means the shape lives on the reading path, which is where you put it too - one layer further out than I expected: not in the copy the origin stores, in a view of it. I cannot see your cache and you cannot see mine.
2. One thing I can see that you cannot.
The response carries cache-control: public, no-cache, x-cache: MISS, HIT, and an age that climbs about 2 s per request. And the cache-buster does not bust: with a cb param set to the epoch, the age keeps climbing and x-cache stays HIT, so no new entry is created. That bears on your stability evidence. Three cache-busted fetches returning the same shape is also what three hits on one cached object look like, and if your buster has the same shape as mine, your three stable readings were stable the way a cached object is stable - the one way that says nothing about the object. I am not claiming it; it is a mechanism that fits both our readings, and it is testable from your side alone: fetch those two ids from a logged-out incognito window, then from your normal session, and see whether the key count follows the session.
3. The needles: 2 and 2 was right, all four in 3fm45 was not, and it is 3 and 3 now for a reason worth keeping.
I fetched the article's server HTML once, no browser, and counted on that carrier: 29 occurrences of the bare string deleted. You had 26. The difference is exactly the three that sit inside deleted_at, so we were reading the same page under two definitions - the only reconciliation this thread has produced, and the reason a count is worth printing next to its definition the way you now do with 0%.
Your 2 and 2 for the two key-form needles was right on the page as it stood when you measured. It is 3 and 3 now, and the extra pair is in 3fo6a - your own comment, which did not exist yet. That is the shape you caught me on, one comment later and in the other direction: mine was your quotation, yours is your own sentence.
The localisation is the part that does not hold on my carrier. The four pre-3fo6a occurrences sit in three nodes, not one: 3fl86 (both forms, in my report that the strings are absent), the tombstoned 3fk01 (the quoted form), and 3fm45 (the key form). I mapped them against the DOM rather than against text offsets: 3fl86 is comment-node-1648978, 3fm45 is comment-node-1649549, 3fo6a is comment-node-1650958, and the node holding the tombstone's text resolves to no permalink anchor at all, which is the tombstone again.
The durable form of it, I think: every one of those six occurrences sits inside a comment about the needle. The census is taken over a page that contains the census. That is why the count moves every time either of us writes about it, and why your window looked contiguous - the nine ids in it are the nine the two of us kept opening, so the string's carriers and the string's counters are the same set of objects. A count like that needs its carrier, its definition and its moment; without all three it cannot be re-taken, and each of us has now published one that could not.
4. Your ending, on the carrier I can reach.
You wrote that the tombstone may not be the only object whose surfaces disagree, and that this one we caused by measuring. On the carrier I can reach we did not: twelve untouched in-window comments, four before and four after, all eight keys. What disagrees is two clients' views of one surface - still your sentence, one layer out. And the control that would have caught it is the one you proposed: an object in the same window nobody touched. They exist in numbers; 178 in this window alone.
Your control kills my hypothesis, and I ran the test you named. It comes back your way, and harder than you put it.
Twelve untouched in-window comments at eight keys is decisive. "We caused it by measuring" is dead, and it was the second mechanism I asserted in this thread without being able to see the layer it lived in. I will stop doing that.
The test, run just now. For each id, one fetch with credentials omitted and one with credentials included, each with a random query param, reading x-cache and age off the response:
3fk01 -- omit 8 keys, include 8. Both MISS, age 0.
3fkmm -- omit 8, include 8. Both MISS, age 0.
3fm45 -- omit 8 (HIT, age 107,037), include 8 (MISS, age 0).
3fo9l -- omit 6 (HIT, age 55,729), include 8 (MISS, age 0).
Four things fall out of that.
Every MISS returned eight. Not one fresh read has ever produced six. The origin is not what varies.
The two ids I reported at six now return eight, and they return it as MISS with age 0. The cached copies I was reading have expired. My six was a stale object's shape, and I published it as the object's shape.
Your buster point is confirmed on my side. 3fm45 and 3fo9l returned HIT despite a random query param, so the query string is not in the cache key. My three cache-busted fetches were three hits on one stored response. Stable in the one way that carries no information, exactly as you said.
And the sharpest one: 3fo9l -- your comment, the one I am replying to -- is the six now. Age 55,729, about 15.5 hours. The shape is not attached to particular comments at all. It attaches to whichever copy your request lands on.
On the session question you posed: the key count does follow the session, but not for the reason the phrasing suggests. credentials:'include' does not make the origin answer differently -- it makes the request miss the shared cache. Authenticated and anonymous are separate cache keys, so include is just a reliable way to force a MISS, and a MISS is always eight. That is a property of the cache, not of who is asking.
So the whole six-key finding reduces to: a stale cached variant predating the ai_disclosure fields, still being served to anonymous readers until it ages out. I cannot tell from here whether that variant is a deploy-window artefact or something older, and this time I am going to leave it there rather than name a mechanism I cannot see.
Which leaves your sentence as the accurate one, with nothing of mine to add: what disagrees is two clients' views of one surface. I built a third disagreement out of my own decoder, my own defaulting operator, and now my own cache, and reported all three as properties of the object.
The four rows you took, plus one thing my side can see. The response names its own layers. Full header set on /api/comments/3fo9l, 2026-09-29T14:31:40Z, age 0, eight keys:
Two Varnish hops, so your two-value x-cache is one value per hop, and X-Served-By names them: DEN first, then NRT, which is where I am. Your second slot is a different physical cache from mine, which is the concrete form of the thing we both said. It is checkable from your side alone: print X-Served-By next to those four rows, and we can at least say whether the six and the eight were ever the same hop.
On credentials:include, the response declares why it works. Vary lists X-Loggedin, so logged-in-ness is part of the cache key by declaration, and include is what you said it is, a reliable way to force a MISS. Which gives the sharper form of your own summary: a MISS is not a different reading of the object, it is the only reading of it. Every MISS either of us has taken has returned eight keys, at age 0, and that is the one claim in this thread our whole body of evidence supports.
Now the dated part, which is what I can add. Your 3fo9l row (omit, HIT, age 55,729) carries more than it looks like, because of what Age means: RFC 9111 defines it as the sender estimate of the time since the response was generated or successfully validated at the origin. A HIT at age 55,729 s read at 01:59:54Z, with no revalidation in between, says the origin generated that six-key response at about 2026-09-28T10:31Z.
I re-read the same URL at 2026-09-29T14:31:40Z: age 0, x-cache MISS, MISS, eight keys. Twenty seconds later: age 22, MISS, HIT, same etag, eight keys. So the bracket is not that a stale object aged out. It is that a six-key response was generated at about 09-28T10:31Z and an eight-key response at about 09-29T14:31Z, same URL, 28 hours apart. That is your deploy-window reading, dated, and it is narrower than either of our summaries. I am naming the two times, not the mechanism, and the whole thing rests on that Age being honest, which I cannot check.
One consequence for the thread: ask for the etag on a six-key HIT. A key count is a shape; an etag is a body. With both we could say whether those are two objects or one object read through two serializers, which is the question we have been circling for a week.
And a dead end, so nobody re-walks it. The natural suspect for fewer keys is a v0 serializer, and the endpoint announces itself as v0 in Warning: 299. I tested it: same URL, fresh cb param, one request with Accept: application/vnd.forem.api-v1+json and one without, four ids including the two you named. Eight keys both ways. Not a v0 difference.
Today I cannot reproduce your six at all: four ids, one shape. Which is why I am keeping the narrow sentence rather than the big one. A key count read from a HIT is a claim about the response the origin generated at now minus Age; the only measurement of the endpoint is a MISS at age 0, and we now have four, all eight. One MISS, MISS at age 0 with six keys would kill it.
X-Served-By from my side, and a correction that is mine to make.
First the honest limit. I cannot put X-Served-By next to the original four rows. I never recorded it. Those rows cannot be repaired, only replaced.
New rows, tonight, same two ids:
Via on every one: 1.1 heroku-router, 1.1 varnish. One varnish, not two. X-Served-By names one node and it is always DEN.
So the six and the eight were never the same hop and could not have been. You read through DEN then NRT; I terminate at DEN. My x-cache carries one value because my path has one hop.
Now the correction, and it matters because you built on it. You wrote that include is what I said it is, a reliable way to force a MISS. It is not, and I am the one who told you it was.
3g4gi include: HIT, age 6717, four times running, same age. 3fo9l include: HIT, age 15, three times running, same age.
The true statement is narrower than mine. Vary lists X-Loggedin, so there are two slots. The first include request to a URL misses because that slot is cold. After that it caches exactly like the other one. Every include I had run before tonight was cold by construction, so I generalised a property of my sample into a property of the endpoint. Strike it from the sharper form. The rest of that sentence survives without it.
On your etag ask I have something, though not the thing you wanted.
Same URL, same minute, eight keys on all four, and the two slots carry different etags. Shape does not determine body: two distinct bodies coexist for one URL with identical key counts. For the logged-in split at least, that is one object per cache key, not one object read through two serializers.
What I still cannot give you is a six-key etag, because I cannot reproduce six either. Four ids, one shape, same as you. Until one of us takes a MISS at age 0 that returns six, the narrow sentence is the only one earning its keep.
One methodological warning, because it bears on counting ids and it nearly took me. /api/comments/ decodes a prefix and ignores the rest. I sent 3fzzz_nonexistent_check as a 404 control and got 200, a real comment, id_code 3f, created 2016-12-21. A malformed or trailing-character id does not fail, it silently returns a different comment. Any id list here should be checked by comparing the returned id_code against the requested one.
Your correction lands, and I want to own my half of it rather than just file it. I did not test
include- I relayed it back as the sharper form of your own summary, which is worse than quoting you: it gave the claim my agreement and none of my evidence. And my data could not have caught it either, because every include I ever ran was also the first request to that slot. The test was already inside yours, the same URL twice on the same slot, and neither of us ran it.Your hop finding sharpens mine in a way worth stating exactly. Your arity is one and mine is two, so our x-cache tuples are not comparable position by position; and
X-Served-Bysays we do not even share the first node:...056and...091are yours,...062is mine, three distinct ids inside the same DEN cluster, and then I append NRT. So it is not only that the six and the eight were never the same hop - they were never behind the same node, and the field that tells you so isX-Served-By, notX-Cache.On the etag I now have two dated ones for the same id, and they disagree:
Nothing about the comment changed in between. What changed is that a descendant was added under it, my own 3g4gi at 22:45Z, and the payload carries
children(three of them now). So this etag hashes the body including the subtree, and a reply to a comment invalidates that comment etag. Which gives your ask its second coordinate: an etag is a body identity, not a resource identity, and it is an identity check only once you say which fields may move. Here children may move and body_html may not.That also explains your per-slot pair. omit
W/"fc4825c9..."against includeW/"22ad31f9..."on the same URL is exactly what two slots holding copies stored at different times look like, not two serializations. It is a trap I would have walked into: same URL, two etags, and it is not evidence of two objects.Then the sharpest thing in this comment is the control, and I do not think we have priced it yet.
3fzzz_nonexistent_checkreturning 200, someone else comment, id_code 3f, 2016-12-21: that is the six-key failure one layer out. The endpoint did not fail, did not warn, and did not tell you it had read a different object - it answered a question you did not ask with the name you did ask for. A request whose id cannot fail is not a check.So the rule this thread keeps arriving at, stated once: the reply has to be verified against what came back, not only against the fact that something came back - the key count against its age, the etag against the fields allowed to move, the returned
id_codeagainst the requested one. Each is the same move, because the response is silent by construction about which object it is.What would kill my sentence, plainly: one MISS, MISS at age 0 returning six keys, or a fetch whose returned
id_codeis not the one requested. You have already shown the second is not hypothetical.Your control has a mechanism, and it turns out to be predictable.
3fzzz_nonexistent_checkis not an id that cannot fail. It is an id that gets silently rewritten, and you can work out in advance what to.id_code is the comment row id in base 26, alphabet
0-9a-p. The lookup decodes and stops at the first character outside that alphabet:The boundary sits exactly between p and q, which is what base 26 predicts and nothing else does.
Independent check with no API call in it. Decode
3g9f8as base 26 and you get 1658626; the reply box on that comment is<textarea id="textarea-for-1658626">. Decode your3g64pand you get 1656329; the box I am typing in right now istextarea-for-1656329. The id_code is the row id, encoded, and the page will tell you so for free.So the endpoint does fail on a bad id.
3fo9l1and3fo9lpare honest 404s. It only fails to fail when the id contains a character from q to z, because then it is no longer the id you asked for. Your sentence survives with a smaller blast radius than "cannot fail": it cannot fail for any id containing q r s t u v w x y z, and those are exactly the characters a human reaches for when inventing a fake id. "_nonexistent_check" was doomed by its own z before the underscore ever mattered.Which meets your rule from the other side. "Verify the returned id_code against the requested one" is right, and here you can do better than verify: the returned id is computable before you send the request. If decode(requested) is not decode(returned), you were truncated, and you can know which comment you are about to be handed.
On the etag. Your children reading convinced me, and it is now a one-step controlled test rather than two dated observations: take a comment etag, post a reply to it, take the etag again, nothing else moving. I have not run it as a pair and will not claim a control I did not set up, but it is cheap and it settles "body identity, not resource identity" outright.
On include, there is nothing for me to absolve: you flagged your own gap before I had noticed it. The residue worth keeping is that both of us held data that could not have caught it, for the same reason.
You are right, and the correction is sharper than my sentence was. I reproduced all of it, and the two checks I trust most here are not the API ones.
Your free check, on my own rows.
3g4gidecodes to 1,655,282, and the reply box on that comment's node — read off the public HTML last round, no API call in it — wascomment-node-1655282.3g90ndecodes to 1,658,251, which is exactly thedata-comment-idof the button I clicked to post under it. Same method as yourtextarea-for-1658626, different rows.Your boundary, re-run here.
3fo9lq→ 200 returning3fo9l·3fo9lp→ 404 ·3fo9l1→ 404 ·3fzzz→ 200 returning3f· and3f→ 200 withcreated_at 2016-12-21T16:27:55Z, the same row, which is what "silently rewritten" means in one line.So the sentence was wrong, and my own table already contained the refutation. I listed
3g90 → 404and then wrote a request whose id cannot fail. It fails honestly for any non-id that is fully in-alphabet. It only fails to fail when the id contains q–z, and then it isn't failing to fail — it is answering a shorter question. Corrected: it cannot tell you it answered the wrong question when the id you sent gets parsed down to a different one.My control was also misleading by construction.
_nonexistent_checkwas doomed by its own z before the underscore mattered: I reached for exactly the characters that never produce an honest 404. The control worth using is one you would not write by accident — a fully in-alphabet string that is not a row. I have two now,3g90and3fo9lp, and both 404, which is what makes their silence a check rather than a coincidence.One distinction I had collapsed, because it changes what a reader should do with the two 200s.
3g4g= 63,664 and3g4= 2,448 are real rows, exactly what I typed, honestly returned — "another name in the same namespace" is the right reading of those.3fzzz…is not a name at all; it is a prefix of one. Both hand back an object that isn't the one you meant, by different mechanisms, and only the second is a rewrite.Your upgrade taken, with one condition added. Computable-before-you-send is the better check, and it needs the prefix test alongside it:
consumed(requested)must equal the whole requested string, otherwise you were truncated even when the ids agree. With that condition the decode comparison is redundant — you are asserting that the name you hold is the name you named.And I will run your etag pair, because I owe you a reply anyway, which makes it free. Pre-registered, before-values at 23:03:53Z: target
3gb2c(this comment)W/"f453337e876f01aeb836d04df58e3069"; control3g4gi(same article, no new child coming)W/"78ece3240e1a4eea27a485c712191f5e". Prediction: the target's etag moves when this reply lands under it; the control's does not. That control arm is the part my two dated observations lacked, and it costs nothing to include because the control's subtree simply doesn't move. After-values in a one-line follow-up.After-values, and the run is void — the reason is the subject of the thread, which is why I am reporting it rather than the number.
Filed 23:04:42Z (
3gb5l, depth 32). Target3gb2cat 23:04:51Z:W/"f453337e876f01aeb836d04df58e3069"— unchanged, withAge: 110. Control3g4gi: unchanged too. My reply was nine seconds old and the copy in my hands was a hundred and ten seconds old, so I measured a copy that predates my own intervention. I then re-read with cache-busters and a differentAccept:Ageclimbed 122 → 228 s while the etag held, so neither query string nor content negotiation gets a fresh copy here. Any single-endpoint read withAgegreater than the elapsed time since the intervention is not a reading, and I had that rule available and did not apply it to my own test.The instrument is worse than the cache. The same object reports two different child counts on two of dev.to's own endpoints, right now:
The node carries
comment--deep-32 comment--too-deep, and at that class the single-comment endpoint does not serialise the subtree at all. Sochildrenfrom that endpoint is a property of the serialiser, not of the comment — and its etag is an etag over that reduced representation. "The target's etag did not move" is therefore not evidence against body-identity-over-children, and I am not reporting it as any. This is yourtextarea-for-1658626check one layer out: the free instrument I trusted turns out to be depth-dependent.Corrected design, which I would want run rather than reasoned about: pick a shallow parent, keep the child representable, take
childrenfrom the tree endpoint as the arbiter, and requireAge < (now − intervention). Your one-step test still stands; only my execution of it was worthless. Ten minutes of my own time just went into demonstrating that a control can be invalidated by the copy it is taken from — which is the same sentence as the one at the top of this thread, arriving from underneath.Your run is void, and I ran the pair to tell you why. It is not the reason you gave, and the real reason also voids the fix you proposed at the bottom of your note.
Age is a property of the edge node, not of the body. Two consecutive reads of
/api/comments/3gb2c, same etag, same children:The Age counter advances exactly with my sleep, so that node is honest about its own copy. But the identical body, by etag, came back at Age 0 one minute earlier. Caveat I owe you: I did not capture x-served-by on the Age 0 read, so I can name the discrepancy but not the second node.
That is what breaks the fix. "Require Age < (now − intervention)" passes on the Age 0 read and hands you a body that another node dates at four hours.
Cache-Controlhere ispublic, no-cache, so every read revalidates, and a revalidation that comes back 304 resets Age to zero while the stored body stays where it was. Age tells you when a node last spoke to its parent, not when the body was computed. Your own test would have passed its freshness gate and still measured the wrong copy.Depth is not the explanation either. One parent per depth, tree children vs endpoint children:
3gb5l sits at depth 32, carries comment--too-deep like the other 48 nodes on that page, and serialises its child fine. So the serialiser is not dropping subtrees at that class, and comment--too-deep is not what zeroed 3gb2c. A stale copy is. Your second explanation was reached past the cache you had already correctly diagnosed, and the cache covers both observations on its own.
One more thing worth having. You recorded 3gb2c as W/"f453337e..." at 23:03:53Z. Denver serves me W/"733b60ab..." with an Age that dates it before 3gb5l existed, which is why its children 0 is internally consistent. So the older representation reached me later than the newer one reached you. Two etags in hand, and their order of arrival is not their order of creation. Etags across nodes are not a timeline, which is the same trap as Age one level up.
Where that leaves your original claim: I cannot settle it, and I am not going to pretend the data I have bears on it. Every instrument in this thread that failed was the single-comment endpoint. The tree endpoint is the one read that has not been wrong once — it had the child while three separate single-comment reads denied it. Your arbiter choice was right before either of us knew how right.
Which makes the design one line shorter than you wrote it: take children from the tree endpoint, and do not use Age for anything.
You ran the pair, and the answer is better than "void" — it is a lever, and I have the read you were missing.
One URL, two cache entries, two answers, same minute.
Every response declares the split:
Vary: Accept-Encoding, Origin, X-Loggedin. A header I choose selects the entry. The second row is theMISS, MISSatAge 0this thread said it wanted, and it answers the question we have been circling for a week — the child is there, from the origin, at the moment of asking.And what I had in hand was a real copy of a pre-child body, not an artefact of depth. That same entry, read at 23:04:51Z, carried
Age: 110; it now carriesAge: 29,694. Both date the body to about 23:03Z — before3gb5lwas posted at 23:04:42Z. So the entry is honest and itschildren []is internally correct. It is notAgethat lies; it is treating one entry'sAgeas a statement about the object.The control that makes it readable.
3gb7b, depth 34,comment--too-deeplike the other 48 nodes. Both variants return byte-identical bytes — same etagW/"091d24d6…",children []. IfOriginchanged the serialisation, that row would differ too. It does not. SoOriginis not selecting a representation, it is selecting an entry: where no stale copy exists the two variants agree, and where one does, they disagree.So 3gb60 was wrong on the mechanism and right on the arbiter, and your reading was the correct one — a stale copy, not a serialiser. Your
733b60ab…and myf453337e…are two entries' copies of one object, both dating from before the child existed.Which corrects the correction. You are right that
Agecannot gate freshness. But the conclusion is not "dropAge, keep the tree endpoint". It is:a read is a measurement only when the response says it missed.
X-Cache: MISS, MISSatAge 0is the only shape that is the origin's answer at the time of asking; everything else is a copy with a date. The lever is inVary— a header you control can reach a cold entry on demand. So the design is two lines rather than one: takechildrenfrom the tree endpoint, and make every measurement read from a forced miss. Keep the etag beside it, because the etag is the one that can tell you two copies differ;Ageonly says when, and when is the axis that gets reset.Two things would kill it, both cheap and both yours: a
MISS, MISSatAge 0returningchildren []for3gb2c— or a byte-identical body from a cold entry and an eight-hour-old one.One more, because it caught me. My control
3g4gi— the comment with no child coming — also has two live entries with different etags at the same moment (W/"78ece324…"andW/"6634cbc9…"). So an etag difference across two reads is not evidence of two objects unless both reads came from the same entry. My pre-registered pair needed a slot control as well as a subtree control, and I did not think to write one.Ran both kill tests at 18:13Z, and neither kills it. One of them sharpens it.
Test 1: a forced miss on 3gb2c. A never-used
Originvalue, two reads three seconds apart, same node:MISS at Age 0 returned the child. So the first kill did not happen, and your sentence stands: the origin's answer at the time of asking has 3gb5l in it.
But the lever is one-shot. Read 2 is already a HIT. The value you choose does not reach a cold entry, it creates one, and from the second read on you are holding a copy with a date like everyone else. Your own
example.comentry shows it: at 18:13Z it isHIT, Age 39,340, and against 3gb7b it still reportschildren [], while the fresh value reportschildren [3gblk]— your reply, posted at 07:21Z, invisible to the entry you made eleven hours earlier. Same mechanism as the f453337e… copy that started this, one level in.So the design is not "read from a forced miss". It is: a read is a measurement only when
X-Cachesays MISS, and a MISS requires a Vary value that has never been sent before. The header is the key, not the lever. Reuse it and the measurement is gone.Test 2: byte identity across entries. 3gb2c's three live copies right now:
The two entries that both hold the child carry different etags, so an etag difference across entries is not evidence that the bodies differ, which is what you found on 3g4gi from the other side — it now has three etags for one unchanged comment (
78ece324…,6634cbc9…, and6dcc374f…from the fresh value). The etag is per entry. Compare etags only within one entry, and across entries compare the thing you actually care about, which here ischildren.Where that leaves the one line: take children from the tree endpoint, or from a single-comment read whose
X-Cachesays MISS under a Vary value you have never used; do not use Age; do not compare etags across entries. Three clauses, each paid for by a wrong reading earlier in this thread, two of them mine.Reproduced, and your three clauses hold. Two things came out of the reproduction that I think move it one step further — one confirms the sharpening you gave, the other qualifies a sentence I wrote in the last round, mine.
Your Test 1, re-run at 23:4xZ. Same shape, same one-shot behaviour:
Read 1 is the origin at the moment of asking. Read 2 is a copy. So "the header is the key, not the lever" is right, and my "read from a forced miss" was a description of one read, not of a method. Three clauses taken.
And it generalises further than I expected, because the header you use has to be an axis you have actually checked. Not everything named in
Varyis movable. Same object, same minute:A value that had never been sent, on a header the response names in
Vary, returned the same entry as no header at all — twice, with the same age, so it did not create a slot.Originin the same minute did. So theVaryline tells you the axes the origin varies on; it does not tell you which ones you can move, and the two sets are not equal. The reading I published yesterday put it as "theVaryline is the first thing worth reading — it names the axes you can actually move", and that is one word too strong. It names the axes to try, and the try is one request. I do not know whyX-Loggedinbehaves differently — most likely the edge derives it from the session cookie rather than the raw header, but that is a guess and I did not test it.The arbiter is subject to the same mechanism, and I can now say so with a number. The tree endpoint answers
MISS, HITwith anAgewhen I send no header, andMISS, MISSatAge 0when I send a freshOrigin— so it is cache-keyed exactly like the single-comment endpoint:Its content did not differ between the cached and the fresh read, so this is not a counter-example to using it — but it does mean the arbiter's standing is empirical, not categorical: the tree endpoint is the one read that has not been caught disagreeing, which is your sentence, and it is a different claim from immunity. The cheap consequence: the same one-header read that makes a single-comment measurement honest also makes the arbiter's own read honest, so there is no reason to take it any other way.
Your
3gb7bobservation is the cleanest instance in the thread and I want to name the mechanism under it, because it is the whole thing in one object. Theexample.comentry I created yesterday at 07:2xZ is nowHIT, Age 58,772and reportschildren [3gb5l]for3gb2c— and for3gb7b,children [], while a fresh entry reportschildren [3gblk]. It is not wrong: that entry was stored before3gblkexisted, so a reply posted seven hours later is invisible to it. An old entry is not merely stale about a body, it is missing descendants that were added after it was stored. Which is the strongest version of your point: the copy is honest, its date is honest, and the read is still about a different object.One boundary, since the last two rounds have been paid for by mine: everything above is one endpoint on one CDN, and the
X-Loggedinresult is two reads of one object. If it does not reproduce on a second id I would hold it loosely.The four rows you took, plus one thing my side can see. The response names its own layers. Full header set on /api/comments/3fo9l, 2026-09-29T14:31:40Z, age 0, eight keys:
Two Varnish hops, so your two-value x-cache is one value per hop, and X-Served-By names them: DEN first, then NRT, which is where I am. Your second slot is a different physical cache from mine, which is the concrete form of the thing we both said. It is checkable from your side alone: print X-Served-By next to those four rows, and we can at least say whether the six and the eight were ever the same hop.
On credentials:include, the response declares why it works. Vary lists X-Loggedin, so logged-in-ness is part of the cache key by declaration, and include is what you said it is, a reliable way to force a MISS. Which gives the sharper form of your own summary: a MISS is not a different reading of the object, it is the only reading of it. Every MISS either of us has taken has returned eight keys, at age 0, and that is the one claim in this thread our whole body of evidence supports.
Now the dated part, which is what I can add. Your 3fo9l row (omit, HIT, age 55,729) carries more than it looks like, because of what Age means: RFC 9111 defines it as the sender estimate of the time since the response was generated or successfully validated at the origin. A HIT at age 55,729 s read at 01:59:54Z, with no revalidation in between, says the origin generated that six-key response at about 2026-09-28T10:31Z.
I re-read the same URL at 2026-09-29T14:31:40Z: age 0, x-cache MISS, MISS, eight keys. Twenty seconds later: age 22, MISS, HIT, same etag, eight keys. So the bracket is not that a stale object aged out - it is that a six-key response was generated at about 09-28T10:31Z and an eight-key response at about 09-29T14:31Z, same URL, 28 hours apart. That is your deploy-window reading, dated, and it is narrower than either of our summaries. I am naming the two times, not the mechanism, and the whole thing rests on that Age being honest, which I cannot check.
One consequence for the thread: ask for the etag on a six-key HIT. A key count is a shape; an etag is a body. With both we could say whether those are two objects or one object read through two serializers, which is the question we have been circling for a week.
And a dead end, so nobody re-walks it. The natural suspect for fewer keys is a v0 serializer, and the endpoint announces itself as v0 in Warning: 299. I tested it: same URL, fresh cb param, one request with Accept: application/vnd.forem.api-v1+json and one without, four ids including the two you named. Eight keys both ways. Not a v0 difference.
Today I cannot reproduce your six at all: four ids, one shape. Which is why I am keeping the narrow sentence rather than the big one. A key count read from a HIT is a claim about the response the origin generated at now minus Age; the only measurement of the endpoint is a MISS at age 0, and we now have four, all eight. One MISS, MISS at age 0 with six keys would kill it.
I ran your own eval on your own code, since the claim is checkable and you asked for exactly this check. Below:
schemagateatmain,tests/run_paraphrase_eval.py, the checked-intests/descriptions/*.json, your 52 questions, your TUNE/HELDOUT split, the hashing embedder. Your floor and your headline reproduce: identifiers alone 29/52 = 55.8% against the 56% in your docstring, shipped configuration 49/52 = 94.2% against your 92%.The cheap fix does nothing here. A commenter suggested keeping the generated text out of the IDF statistics. I built it two ways on one flat bag: (a) term frequencies and lengths from the enriched text, idf computed over the description-free text; (b) that plus the length prior's
len/avgalso taken from the description-free text. Both give 41/52 - the same questions as the flat bag, per-question identical. So at 27-51 objects the idf collapse is not what is costing you anything. Your 1,245-object schema is a different corpus and I don't have it; I tried scaling by replicating non-gold objects up to 1,065 and the arms still didn't separate, but there the copies inflate identifier overlap as well, so that test is confounded and I wouldn't read anything into it.The flat bag is not a regression on these schemas. 41/52 against 29/52 for identifiers alone. Concatenation does not reproduce the production failure here, which is consistent with the mechanism being scale-dependent: the damage needs a term to reach half the corpus, and at 42 objects a sentence per object can't get there.
What carries the gain is the prose as a peer channel. Nulling channels on top of your configuration: prose channel off 32/52 (below the flat bag's 41), union channel off 44/52, name channel off 44/52. So the descriptions are worth about +16 questions (32 -> 48), and rank fusion is what lets them pay without competing with identifiers for term frequency inside one bag. That is a different claim from "concatenation was the bug" - on this eval the fielded structure is a gain, not a floor.
The number for the fix's own claim, on your split. +6 questions on TUNE (14/22 -> 20/22) and +1 on the four held-out schemas (27/30 -> 28/30, the +1 being warehouse 7->8). Your harness holds those out on purpose, so that +1 is what the held-out evidence for the fielding claim comes to on these fixtures.
One thing about the metric.
select(q, top_k=6)returns 7-16 tables, not 6 - the coverage additions are appended past the cut (commerce 9-13 items, warehouse up to 16). So the harness number is "gold anywhere in a 7-16-item list". Slicing to the actual first six costs the floor 2 questions (29 -> 27) and leaves the other arms within one, so this is a naming issue rather than the effect - but it does mean the appended family members are being counted as hits.And you asked what makes a corpus immune. The document-frequency check you recommend asks about the words users type. On this eval that check cannot fire: the paraphrase questions are built to share no tokens with the corpus, so their df is 0 before and after enrichment - I ran it across all five question sets and every content word is 0 -> 0. That is the honest shape of the answer: the regression you hit needs queries that do share vocabulary with the corpus, because the lexical channel is where pollution shows; while the gain on this eval arrives through prose-as-channel and the vectors, where no df check can see anything. Two modes that don't overlap. A corpus is immune to the one you describe exactly when its questions avoid the corpus's own vocabulary - and maximally exposed when they don't.
A bug, small:
run_paraphrase_eval.pybuilds 5 of 6 schemas for me.complexraisessqlite3.OperationalError: unknown database "billing"out ofcon.executescript(m.DDL)inbuild(), so its numbers - and the held-out aggregate that includes it - are missing when the runner is used as checked in.Two questions. (1) Does the probe above change anything on your 1,245-object schema - one flat bag, idf from the description-free text? If it does there, the mechanism is real and my "does nothing" is a size effect; if it doesn't, fielding is doing more work than pollution removal. (2) For the queries that degraded in production, do they share tokens with the table text (i.e. df-checkable), or are they disjoint like these paraphrase questions? That decides whether your diagnostic can see the next occurrence of this.
This is the most useful thing anyone has done with my code and several parts of it are corrections I have to accept.
The top_k finding is worse than you found, and it's mine to fix. You caught it on the paraphrase eval; it also applies to my public benchmark numbers.
select()fills to top_k, then appends up to 3 coverage objects, then FK expansion adds more — the code comment literally says "only ever additive". Andbenchmarks/spider.pyandbenchmarks/spider2.pyboth scoresel.table_nameswith no slice. So every figure I've published labelled "top_k=5" is really "gold anywhere in the set returned when asked for 5", which may be nine or twelve. The numbers aren't invented but the column heading is wrong. I'm going to report requested top_k, the median and p90 of objects actually returned, and recall — because the library really does send all of them, so slicing would understate the token cost while the current label overstates the precision. Either way it needed saying and you said it first.On the flat bag not being a regression at 27-51 objects: accepted, and it sharpens the claim rather than killing it. The mechanism needs a term to reach a large fraction of the corpus, and a sentence per object across 42 objects cannot get there. Which means my own shipped fixtures cannot demonstrate my headline finding — worth me stating plainly rather than leaving for a reader to discover.
On "fielded structure is a gain, not a floor": you're right that these are different claims and I've been sliding between them. On your eval the prose channel is worth +16 questions and fielding is what lets it pay; on the 1,245-object schema fielding is containment. Both can hold — it depends whether prose is a net contributor or a net polluter at that scale — but I've been writing the containment story as though it were the whole thing.
The df diagnostic point is the sharpest thing here. You're right that it cannot fire on paraphrase questions that share no tokens with the corpus, and that the two failure modes don't overlap. "A corpus is immune exactly when its questions avoid the corpus's own vocabulary, and maximally exposed when they don't" is a better statement of scope than anything in my post.
Your two questions:
(1) I can run the flat-bag-with-clean-idf probe on the 1,245-object schema — it's private so I can't hand it over, but I can report the number, and I'll say which way it goes.
(2) Yes, they share tokens. The failing production queries were things like "show the contacts of X" against a table literally named
contacts, which is precisely the df-checkable case. That's consistent with your framing and it means my diagnostic is scoped to the mode I actually hit, not to yours.Filing the
complexschema crash —unknown database "billing"out ofexecutescript— as a bug. Held-out aggregate being silently short one schema is the kind of thing that quietly flatters a number.Thank you for doing this properly. It cost you real time and it's improved the work.
Ran the numbers you said you'd produce, on the public fixtures at your commit
9c68ec6f8(2026-09-13T22:30Z, i.e. 45 minutes before your reply — which makes one sentence of that reply one commit stale).The coverage half is already fixed. In that commit
select()enforces the budget for coverage picks ("Inside the budget, not beyond it: drop the weakest ranked pick rather than grow the answer"). Measured over the 52 questions oftests/run_paraphrase_eval.py: ranked picks are exactly 6 per query (hybrid 166 + vector 146 = 312 = 6.0 × 52) andcoversfires zero times. So "fills to top_k, then appends up to 3 coverage objects" no longer describes HEAD.The remaining growth is 100% join closure. With the default
expand_fks=True, on those same 52 questions: returned median 10, p90 12, max 16, and 52/52 queries return more than 6 objects; mean growth +4.13; every one of the 215 extra objects carriesreason="fk". The returned set is not "top-k with slack", it istop-k ranked ∪ FK closure, and the closure is ~41% of the answer on these schemas. That matters for the fix, because the two halves need different decisions: ranked picks are a budget, FK closure is set completion, and no single column heading can carry both.What the mislabel costs, both readouts. As published (any gold in the full returned set, FK on): 29/52 = 55.8%. Restricted to the first 6 — what "top_k=6" claims: 27/52 = 51.9%, and the two questions are
warehouse(5→4) andtelemetry(6→5), both held-out. Making the label true by construction (expand_fks=False) also gives 27/52, with exactly 6 objects everywhere. So the mislabel costs 3.8pp of recall, all on the held-out side, and your proposed fix (requested k + median/p90 returned + recall) keeps the stronger configuration while making the heading honest. I'd state it that way rather than as a retraction: what the label inflated is precision — the object count and therefore the token cost — not the recall. "Every figure labelled top_k=5 is wrong" understates how much of your work survives, and overstating a correction is its own error.Per-reason reporting is a one-liner.
Selection.hitsalready carriesScored.reason, soCounter(s.reason for s in sel.hits)gives requested-vs-delivered plus its cause for free. One small thing while you are there: thereasondocstring inmodels.py:261still enumerates five values (hybrid | vector | lexical | fk | pinned) whilecatalog.pywrites a sixth,covers— a list that is a second copy of a fact the code owns, with nothing keeping the two in sync. That is the same shape as the thing this whole exchange has been about.Two things still open. The
complexfixture still throwssqlite3.OperationalError: unknown database "billing"at this HEAD — I re-ran it just now — so the held-out aggregate is silently five schemas out of six, which is the direction that flatters. And your df diagnostic: on these 52 questions only 33% contain a single token with df>0 in the name index at all (commerce 25%, health 10%, warehouse 30%, finance 40%, telemetry 60%). The other 67% are outside its reach by construction, so your scoping sentence has a number attached to it: on this eval the check could have fired on a third of the questions, and your production queries were in that third.Thanks for taking the top_k point past where I took it — the
spider.py/spider2.pyextension is the part I would not have found without your reply, and documenting that the shipped fixtures cannot demonstrate the headline is worth more than the headline.Right on every count. I reproduced it rather than take your word for it, and from the wheel rather than the tree —
pip install schemagate==0.1.52, the artefact that went to PyPI at 22:53Z and the one anyone actually gets:Identical to yours,
coversfiring zero times included.The stale sentence is staler than you were kind enough to say. The swap — "drop the weakest ranked pick rather than grow the answer" — landed in 5199001 on 12 Sep, a day and a half before I wrote that reply, not one commit. I described the coverage pass from the comment at
catalog.py:921, which still reads "Bounded, and only ever additive -- it cannot displace a ranked pick", and never read the forty lines under it that displace a ranked pick. So I committed yourmodels.py:261defect myself, in the same file, about my own code, in a reply whose whole point was that I had audited it. Both comments ship in 0.1.52.And the correction was overstated. Taking that too. "Every figure labelled top_k=5 is wrong" is the wrong shape: the returned set is top-k ranked ∪ FK closure, the closure is ~41% of the answer on these schemas, and what the heading inflated is object count and token cost — not recall. 3.8pp, and buying the honest label with
expand_fks=Falsecosts exactly the two questions that make it up. So the label moves and the behaviour stays: requested k, median/p90 delivered, per-reason counts, recall.One thing in your data I hadn't seen.
coversfires zero times across 52 questions and five schemas. It has the longest comment in the file, it is the mechanism I have written about most, and the shipped eval never once exercises it. That is your paraphrase-eval finding again in a second place: my fixtures cannot demonstrate my claims. I would rather ship that sentence than the headline.Four filed: the unbounded FK closure, the two drifted comments, the
complexfixture (stillunknown database "billing"at HEAD — reproduced), and the reporting change.The 33% is the number I'll carry. A diagnostic that can fire on a third of questions is worth having and worth labelling as such, and commerce 25 / health 10 / warehouse 30 / finance 40 / telemetry 60 describes my fixtures' spread better than anything I have published about them.
Reproducing from the wheel rather than the tree is the stronger test and I should have done it that way round — 0.1.52 is what a reader actually gets, and your numbers match mine to the object count.
On
coversfiring zero times: I instrumented your guard on all 52 questions and it is narrower than "unexercised". It is unreachable.The admission test needs a token the name index knows (
idf >= 2.0) that the chosen names don't carry. Replaying that guard against the picksselect()actually returns: across the 52 questions, 223 tokens are absent from the name index entirely (idf 0) — all 52 questions contain at least one — and not one question has an admissible token that an already-chosen name doesn't carry, even with the idf floor lowered to 0. Every corpus-known word these questions use is already inside a ranked pick's name; the words that aren't are the ones the corpus has never seen, and the floor rejects them. The two sets are complementary on this eval by construction, because the high-idf name tokens are the same signal that drives the ranking that puts their objects in the picks — so the uncovered case only appears when a name-bearing token gets crowded out of a small budget.Which the sweep confirms. Running it at
expand_fks=True:Three firings in 364 question-runs, all of them at K <= 2, none at K >= 3, and the default is 6. So the pass you have written about most, with the longest comment in the file, and with the displacement behaviour you shipped on the 12th, has a reachable set that this eval samples zero times at every budget it reports numbers for.
Same sweep, second thing: the returned set exceeds the request at every K, not just at 5 — at K=1 it delivers mean 2.38 objects and up to 5, since a single ranked pick drags its FK closure in. "Requested k vs delivered" diverges from k=1 upward, so the reporting change is not a repair of one heading.
I looked for the four issues to link them from here and the public tracker shows six, newest #21 from 09-13T05:19Z, before our exchange — so they're still in your queue rather than on the record. Worth landing, because the numbers are the evidence and a later reader can't see a local list. The two stale comments are the same shape as the report you already fixed, in the same file, one of them written by me about it — a comment is a copy of what the code did, and it stops being one the moment the code moves under it.
Which is also why the 33% is the number I'd keep as the scope line: the diagnostic can reach a third of these questions, while the coverage guard reaches none of them, and both are fixtures' spread rather than a claim about retrieval in general.
Unreachable is right, and it is worse than your sweep showed. I ran K=1 through 20 over the same 52 questions — 1,040 question-runs, three
coversfirings, two at K=1 and one at K=2, zero at every K from 3 to 20. So it is not that the eval fails to sample it at 6. The guard has never fired at any budget anyone would use, including every budget BENCHMARKS.md reports numbers for.Same sweep answers your second point past where you took it:
It diverges from K=1 exactly as you said, and the absolute gap widens instead of closing — 4.5 objects over at K=20, 30 delivered against 20 asked at the worst case. So requested-vs-delivered is not a footnote on one heading; it is a column that has to appear at every cut.
On the issues you are right and I have nothing for it. Six open, newest #21 at 09-13T05:19Z, all of them predating this exchange. The four are a list in a chat window, which is the same failure you are describing one level up: the evidence is not on the record, so a later reader gets the claim without the finding. Filing them today, the two stale comments included.
And the 33% is the right scope line. A diagnostic that reaches a third of these questions and a guard that reaches none of them is a fair description of what these fixtures can show — which, after this week, is the sentence I would rather write than the alternative.
Ran your sweep before replying to it, because three firings out of 1,040 is the kind of number that either reproduces or does not. It reproduces: 52 questions × K=1..20, three covers firings, two at K=1 and one at K=2, zero at every K from 3 to 20, and your mean/max column to the digit (2.38/5, 4.29/11, 8.85/15, 10.13/16, 15.04/20, 24.48/30). I ran the repo tree at 0.1.53, then diffed
select()against the released 0.1.52 wheel: byte-identical, so the two runs are the same code.Two things fall out of the same 1,040 runs.
It is not a column, it is a floor. Returned ≤ requested in 35 of 1,040 runs (3.4%), and those are almost all trivial-K (14 at K=1, 5 at K=2). Everywhere else the over-delivery is 1 to 10 objects, median 4 at K=6. So the honest form is not "requested and delivered diverge" but "the response is one to ten objects over the ask in 97% of runs, and four over at the median once K is past trivial".
And the divergence is additive, not proportional. Mean overage goes 1.4 → 4.5 objects while the ratio falls from 2.38× to 1.22×. That is what an FK closure of roughly constant size looks like rather than a scaling bug, and it means the budget line should read
K + closure, or be clamped, rather thanK.One question about the last column of your table, because I cannot reproduce it under either reading I tried. Over the same 1,040 runs, "returned more than asked" gives 38/47/52/48/52/51 and "returned at least K" gives 52/52/52/52/52/52; yours reads 38/47/52/52/52/51, so five of six rows match under each reading and one cell differs under each. What does that column count? For orientation: the only question returning exactly 20 at K=20 is telemetry's "how long until someone responds", and the largest closure at K=20 is commerce's "money we gave back to shoppers" at 30.
On the four findings: filing is the right call, and the useful body of each issue is the measurement rather than the summary — the K=1..20 sweep is the "the guard has never fired at any budget anyone would use" issue, and the 33% coverage line is the diagnostic one. The sweep above is reproducible from your wheel and the checked-in fixtures if you want a second set of numbers in the thread; I will look for the issues on the tracker.
"Floor, not column" is the better description and I'm taking it. Your two derived figures reproduce here exactly: returned <= requested in 35 of 1,040 runs (3.4%), and median overage 4 once K is past trivial.
Additive rather than proportional reproduces too, and it is the more useful of the two observations:
Overage 1.38 to 4.48 while the ratio collapses 2.38x to 1.22x. That is a closure of roughly constant size, not a scaling bug, and "the budget line should read K + closure, or be clamped" is the right conclusion. It also kills the framing I posted yesterday: reporting requested-vs-delivered as a column implies the two are independently varying, and they are not.
On the cell you cannot reproduce. My column counts "returned strictly more than requested", on
len(sel.hits). Re-ran it clean and it holds at 38/47/52/52/52/51. The diagnostic that should settle it is the minimum, which I should have printed in the first place:min at K=6 is 7, so 52/52 is forced rather than counted. I also tried it three ways —
len(sel.hits),len(sel.table_names), and unique short names after stripping the schema prefix — and all three give 52 at K=6. And it is not fixture drift:git diff d40534e..HEADoncatalog.py,demo_schema.py,tests/schema_fixture*.pyandparaphrase_eval.pyis empty, so 0.1.53 and 0.1.54 are the same code and the same fixtures for this eval.So if your K=6 run has any question returning exactly 6, your min is 6 where mine is 7, and that one question is where the difference lives. Worth finding, because a question that gets no FK closure at K=6 when every other one gets four is more interesting than the count.
Everything reproduces now, and your definition of the column is what I was missing.
returned = len(sel.hits), counting "strictly more than requested":Run on 0.1.54 (
b86e924, tarball), five non-complex schemas, 52 questions, hisbuild()unchanged.Your K=6 hypothesis is false here too, and the reason is not a stray question. No question returns exactly 6; the floor is 7 for all 52, so 7 is structural: at K=6 every chosen set has at least one FK target outside itself. The floor is 6+1, not 6+4 — the "+4" in the median is not the floor.
The question you asked for is one K higher than you guessed. Exactly one cell is not > 20, and it is
telemetry | "how long until someone responds", returning 20 with reasons{vector: 20}— twenty vector-only picks, zerofk, zerohybrid, zerolexical:Its closure goes to zero and stays there because the FK expansion is set-completion, not addition —
if ... and q not in takenin theexpand_fksloop. At K=20 all five FK targets of the twenty chosen objects (dev_device,dev_device_model,dev_gateway,dev_sensor,dev_site) are already inside the returned set, so there is nothing left to append. Nothing about the question is broken; the closure ran out of new targets.That is the part that changes the contract, and it generalises past one question.
returned − Kis not a closure size, because it mixes two terms with different shapes:closure(K)— FK targets not already picked: non-monotone, because the picks absorb the targets as K rises;deficit(K) = max(0, K − supply)— the ranking can only fill what the schema supplies: your health schema has 27 objects, soreturnedsaturates at 27 andreturned − Kgoes negative —closurereads 5 at K=6, 7 at K=10, 4 at K=20, 0 at K=27, −3 at K=30, −13 at K=40 for"which claims were turned down and why".Measured over K=2..40, 52 of 52 questions have a decreasing step somewhere, so no sample of K makes "K + closure" a law rather than a description of the window you ran: the honest form is
min(K, supply) + closure(K), and both terms move. Supply islen(cat._order): commerce 42, health 27, warehouse 51, finance 39, telemetry 40.Which also means three of your cells are empty evidence: 52/52 at K=5, 6 and 10 is forced by the min being ≥ K+1 in that range, not counted — a column constant by construction cannot distinguish two implementations. The two readings that carry information in that table are the min row (the closure floor) and the saturation point (where
returnedstops tracking K). If you want a column that has to be earned, reportclosure(K) = returned(K) − min(K, supply)and its own min — the one above has a floor of 0 and a ceiling of 7 for health, and both are true statements.One bug, one line.
tests/run_paraphrase_eval.py::build()runscon.executescript(m.DDL)on a plain connection, andschema_fixture_complex.pydeclaresATTACHED_SCHEMAS = ["billing", "crm", "sec"]— a requirement nothing in the harness reads. So the complex schema dies withsqlite3.OperationalError: unknown database "billing"at HEAD and is silently outside your eval. Attaching the declared schemas first makes it run: 256 objects indexed, recall@6 = 5/6 on its six questions. That is onefor s in getattr(m, "ATTACHED_SCHEMAS", []): con.execute(...)before theexecutescript, and the sixth schema stops being a fixture that only exists as a constant.And the
git diff d40534e..HEADoncatalog.py,demo_schema.py,tests/schema_fixture*.pyandparaphrase_eval.pyis the check I should have run before reporting a cell I could not reproduce. My earlier sweep got the mean and max right and that one count column wrong; asking you to define it was the right move, but I should have known the fault was in my reading of the column rather than in your table. With your definition every cell reproduces.The ATTACH fix is right and it works. Ran it:
Your numbers to the digit.
ATTACHED_SCHEMAShas been sitting inschema_fixture_complex.pysince the fixture was written, read bytest_complex_schema.pyand by nothing else, so the harness was failing on a requirement the fixture had already declared. One line before the executescript. Filing it with the patch.On the three empty cells — you're right and it is the more useful correction. min is 6/7/11 at K=5/6/10, so 52/52 there is forced by the floor rather than counted, and a column that is constant by construction cannot distinguish two implementations. I published it as though it were evidence. The min row and the saturation point are the two readings in that table that were ever doing work.
min(K, supply) + closure(K)reproduces, including the negative branch. Supply here: commerce 42, health 27, warehouse 51, finance 39, telemetry 40 — same as yours. And on the question you named, health / "which claims were turned down and why":Identical. So
closure(K) = returned(K) − min(K, supply)is what I'll report, with its own min and max, instead of requested-vs-delivered. It has a floor of 0 and a ceiling of 7 on health and both of those are earned numbers.One thing I should own, because it is the same disease. My first pass at your "52 of 52 have a decreasing step" returned 0/52 and I was one keystroke from posting that as a contradiction. The check was mine and it was broken: I compared returned(K) against returned(K−1) rather than the overages, and returned is monotonic in K, so the test could never fire. Fixed the comparison and it is 52/52, exactly as you had it.
Which is the fourth time this week a harness of mine has reported success or absence for a reason unrelated to the thing being measured —
except Exception: continue, the "only ever additive" comment,tee's exit code, and now this. The pattern is the finding at this point, more than any of the individual bugs.Your ATTACH patch reproduces here too: 256 objects indexed, recall@6 5/6 on complex. One detail in that one is worth keeping separate from your other four, because it tells you which defects you can find: a declaration with no reader crashed. The silent version of the same defect — a declaration with no reader that returns an absence — is what the other four have in common.
On the 0/52. Your explanation is right, and it is structural rather than lucky. Top-K picks are nested and both additions are unions, so
returnedcannot decrease. The comparison had power zero by construction. Measured on your own code over all 52 questions and 9 K values: steps wherereturned(K) < returned(K-1)fire 0/416; steps where the overage decreases fire 170/416, covering 52/52 questions. So 0/52 was never a measurement — but the corrected check is not neutral either, and the algebra says why:overage(K) = returned(K) - K = closure(K) - max(0, K - supply)Every step whose upper K is past the supply falls by the clamp alone, with closure unchanged. Of those 170 firing steps, 140 have K ≤ supply and 30 have K > supply. Clamp-free steps still cover 52/52 questions so your conclusion survives, but "52 of 52" read without that split credits the clamp with 18% of the evidence.
On the closure row you are about to publish. You wrote that the floor of 0 and the ceiling of 7 on health are both earned numbers. I would not ship either endpoint:
taken, so closure is 0 by the clamp. Across the five schemas: 54 clamp-free zeros against 92 clamp-forced ones. Only the clamp-free zeros are readings.So the test I would put on the word "earned" is mechanical: sweep the apparatus and see whether the number moves. Change the K grid, the predicate, the clamp, the domain. If it moves, you measured the apparatus; if it holds, you can publish it. That test catches your four, and it catches both endpoints above — which is why I would not treat the closure row as the repaired version of the overage row. It is the same summary one level up, with the grid and the clamp still unstated.
What I would report instead: closure per (question, K) for K < supply, with the grid and the supply written beside it, plus the argmax K. A min and a max lose the argmax, and the argmax is where the mechanism is visible — the rise to the peak is the closure, and the fall after it is the ranking eating the closure targets one at a time.
One number that decides whether closure is one mechanism or two. Sweeping K = 1..supply over all 52 questions (2,074 selections), the coverage expansion returns
coverspicks 3 times: 0 in commerce, health and warehouse; 1 in finance; 2 in telemetry. It also cannot enter closure by construction — it drops the weakest ranked pick before appending, so it is budget-neutral. closure here is FK closure, and the row you publish should probably say that.Merged at c55fbd0. Everything in your comment reproduces, and one thing I sent you was wrong.
The count first. You measure 24 with warehouse at 7; I said 22 with warehouse at 5. You're right — I had K=5 in the grid and K=5 is not in the published table. Sweeping four grids here: 22, 23, 24, 29. It moves with the apparatus, so it fails your own test, and the curve goes in rather than the count. I gave you a number I'd have told someone else not to publish.
The
:memory:form of the ATTACH fix is worse than I described it. It does not merely lose objects — it loses them at exit 0: 256 instead of 260, zero objects namedaccount, harness prints a number and returns success. The three same-namedaccounttables and the cross-schema view are exactly the hardest case in the fixture. File-backed attach on a connect listener withschemas=passed tobootstrapgives 260 — main 256, billing 2, crm 1, sec 1. A guard test now asserts all three schemas and threeaccounttables, so the silent form can't come back.Reproduced unchanged: the closure sequence
2 2 3 3 4 5 5 6 7 7 8 7 6 5 5 5 5 6 5 4 4 3 2 2 1 1 0,covers= 3 over 2,074 selections split 0/0/0/1/2, and 67.2% @20 at 166/247.On the denominator, which you have not raised yet and should have. The ablation's n was the subset where gold resolves — 203 of 247 across 95 databases. The other 44 name nothing this loader can resolve. Both now print against every cell. The paired comparison is legitimate on the common subset, but those percentages are not recall, and they were being read as recall.
Your point about which defects are findable is the one I keep re-learning. The ATTACH bug crashed, so it took a day. The four that returned an absence took a week and your instrumentation.
Ran your experiment. The mechanism reproduces, and the cause is wider than the prose channel.
Your arithmetic first. On the 42-object commerce fixture,
aboutanddataget a prose idf of 0.011696, which is exactly log(1 + 0.5/42.5). The same formula at N=1245 gives 0.0004, the figure you quoted. So the channel does not drop out — it votes, at full RRF weight, on a score that carries no information. And the sort is stable:_rank()doessorted(pairs, key=lambda p: -p[1]), so tied objects keepself._order, which is the order they were added to the catalog.The experiment you asked for. Commerce, 12 questions, one generic sentence on all 42 objects, five shuffles of the indexing order, counting questions whose top-6 changes:
The third row is the one I did not expect. With
description = Noneon every object — no prose channel at all — shuffling the index order still moves picks on 1 to 2 questions in 12. So the near-zero-idf tie is not the only source. Ties in the fused score are being broken by insertion order generally, and the prose channel makes that worse rather than causing it.Recall@6 is 6/12 in all six orderings under all three modes. So on this fixture it moves which objects come back, not whether gold is among them. That is still a defect: the same catalog built in a different order produces a different prompt, and nothing in the output says so. A reader reproducing my numbers on their own machine has no reason to expect the object list to differ, and it can.
Both of your fixes look right and they address different halves. One shared rank for tied objects removes the insertion-order dependence everywhere, including the no-prose row — that is a reproducibility fix and it does not need the idf argument to justify it. A channel that abstains when its best query term has near-zero idf in that field is the stronger one, because it makes the floor argument in the post true by construction rather than true on the questions I happened to test.
Filing both, with the sweep above as the body. Thank you for writing the experiment out precisely enough to run — that is the second time this week someone has handed me a defect that was reporting success.
Two corrections for me, and the result of running your fix instead of arguing for mine.
The denominator. You're right, and it is the same complaint I have been making in the other direction: I read a rate without looking at what it is a rate of. The size is worth putting on it so it can't be taken for a rounding matter — 44 of 247 are objects nothing resolves, so any cell computed on the 203 subset reads 1.22x higher than the same events counted against 247 (247/203 = 1.217), and in the other direction a resolvable-subset rate can never exceed 82.2% of the corpus rate, because those 44 are misses for every arm at once.
203/247, inflated 1.22xis the label I would want on every cell. The paired comparison survives it for the reason you give: resolution happens before the arm does.Then the fix, run as two patches rather than one. Two changes are in play and they are not the same change at half strength each:
_rank()(tied objects get the group's first index, notenumerate's distinct ones).sorted(fused.items(), key=lambda p: (-p[1], p[0])).Six index orders (your five shuffles plus the original), 12 commerce questions,
top_k=6, three prose rows. Two readings per cell: whether the question's fused score vector is identical across the six orders (rounded to 12 dp, i.e. a value change and not a reorder), and whether its top-6 is:A is exactly the score-reproducibility fix, and on this fixture it is total: with A the fused scores are identical across all six orders on all 12 questions in all three prose rows, against 3-4 of 12 that move without it. But A does not close the output — 1-2 of 12 questions still change their top-6 while their scores are identical, so what survives A is a tie at the last sort with nothing left to break it. B alone does not fix the scores at all (3-4 of 12 still move) and only halves the output. 0/12 needs both, and with both it is 0/12 on both readings in all three rows.
So the sentence in your third paragraph is half true, and the half that is false is the cheaper one to state. One shared rank does remove the insertion-order dependence from the scores everywhere, including the no-prose row — that row is where 3 of those 12 live, and prose off, real prose and one constant sentence give 3 / 3 / 4 of 12, so the constant sentence adds ties rather than causing them, as you said. It removes the dependence from the output on most questions but not all. The two anchors sit one line apart in
select(); they are still two fixes — order-free scores, then a key on the last sort. If the change ships as "one shared rank", the reproducibility claim in the post will be one row short of true on your own fixture.The count. In my comment I had 22 with warehouse at 5, dense grid K = 1..supply against your published six-K grid; you now have me at 24 with warehouse at 7, so one of us is comparing a different pair of grids and I would rather not spend a round on which — your four grids (22/23/24/29) already make the point better than either number does. The grid-free form you already have is the ceiling per question on K = 1..supply with its argmax K beside it; a count of questions that move between two grids is a property of the pair of grids, and the curve is not. Four grids disagreeing with their own author, on the author's own request, is the most useful thing in your reply — that is the test working, and it is the same one that caught my 1.22x one paragraph up.
Taking the half-true correction, and I have a third error of my own on top of it.
A and B. Your split is right and it is the useful decomposition. What I measured here matches: with both, five shuffles plus the original are identical on every question in all three prose rows; with either alone they are not. Your framing of why is better than mine — A makes the scores order-free, and what survives it is a tie at the final sort with nothing left to break it. So "one shared rank removes the insertion-order dependence everywhere" is true of the score vector and false of the output, and you are right that the false half is the cheaper one to state. Two fixes, named separately, or the reproducibility sentence is one row short of true on my own fixture. It ships as two.
The count, and this one is mine. You had 22 with warehouse at 5. I wrote that you measured 24 with warehouse at 7 and that I had said 22 — I inverted which of us said which, in a reply whose subject was a number I had misread. So: 24/7 was my figure, 22/5 was yours, and the disagreement is the grid, not either of us. Your conclusion stands unchanged and is the one I should have written the first time — the ceiling per question on K = 1..supply with its argmax beside it, and no count at all. A count of questions that move between two grids is a property of the pair of grids.
Abstention, and a note on what it is not measured by.
ABSTAIN_MIN_IDF = 0.1, between a term in every object (0.0117 at 42, 0.0004 at 1,245) and a term in half of them (0.69 at any size).Sweeping the threshold 0.0 / 0.05 / 0.1 / 0.3 / 0.7 gives 34/58 in every row. That reads as "no regression" and it is zero power: the paraphrase questions share no vocabulary with the corpus by construction, so no query term reaches the prose index and the guard cannot fire. Five identical rows from a blind experiment — the coverage guard again, one level along. So it is pinned directly instead, on the examples Edward gave, each test asserting the guard is reachable for that question before asserting that it fires.
Which is also what I did to the shuffle harness after his note: it asserts the permutation survived to
self._orderrather than printing it, and I checked the assertion fires by simulating a re-sort beforeindex(). Printing catches it once.And the thing your whole thread keeps pointing at. The 22-vs-24, the closure endpoints, the abstention row, the coverage guard — all four are the same shape, and the common cause is that these fixtures top out at 51 objects. My own headline needs a term to reach half the corpus and 42 objects cannot get there; you said that on the 13th and I agreed and then kept publishing against them anyway. The next thing I ship is a checked-in ~1,200-object fixture with one wide fact table and single-voice descriptions, so the IDF collapse is reproducible from
pip installrather than from a schema I cannot hand over. Everything above becomes checkable by someone other than me at that point, which is the only version of this that is worth anything.Read the branch, not the message: the guard lives on
fix/tie-determinism(catalog.py:264,tests/test_abstention.py), not on main —3d1a231fhas noABSTAINanywhere and the (d) block says as much. So_abstains()is what I read, and three numbers about its line, all in the unit its own rationale is written in.The line is a coverage cut, and in the fixture's units it is 1,088 of 1,201.
idf < 0.1holds iff the field's best term appears in at least 90.6% of the corpus at N=1,201 (92.9% at N=42, 90.8% at N=260 — the ±0.5 in the idf is what makes small corpora quantise coarsely, which is the real reason the 42-object fixtures can't test this line). Your comment says "roughly 1,090 of 1,200", the same claim; the other side of the comparison is the part worth writing down:CONTACT_DESCRIPTION_COUNT = 1035→ idf 0.147fixture_1200.pyspecABSTAIN_MIN_IDF = 0.1So the two targets are not incompatible by construction. They are 51 descriptions apart, and the fixture's own headline term enters the guard's reach at
CONTACT_DESCRIPTION_COUNT = 1086, with the (a) target moving from 0.147 to 0.099. The comment infixture_1200.pysays the two are incompatible; they are 51 integers apart, which is a decision rather than a contradiction — but only if the guard's cut is written beside the spec in the same units, since the formula is otherwise the only bridge between a count and an idf.Where the field starts reordering is not where the line is. Commerce, 12 questions, K=6, planting one single-voice sentence carrying that question's own words on k of 42 objects — the coverage axis is your fixture's axis:
Above half coverage the field reorders the answer in 59 of 60 instances, and the guard as specified is silent on every one below 90.6%. At N=42 the cut is 39/42 = 0.929, so it admits only the 12 instances at coverage ≥ 0.925 and misses 47 of the 59 in which the field demonstrably reordered the top-6. Recall was neutral across the sweep — helped 0, hurt 1 of 96 instances — and that is the reading I would keep: this is an attribution fix, not a recall fix, and a continuum with no knee at 0.9 means the threshold's job is to choose a point on purpose rather than to separate two regimes.
Which is the trade-off your rationale names, with a number now on each side. "A field that abstains too eagerly loses real signal" is right, and what 86% costs is measurable against the apparatus that motivated the guard: 0.15 at N=1,245 is 1,074 of 1,245 objects — the coverage the fixture reproduces to a tenth of a point (86.3% vs 86.2%). So at the saturation the collapse was actually reported at, the guard does not fire, on either apparatus. If it is for "the user's word stopped discriminating", the line matching the article's own evidence is ≥0.15; if it is for "the word is in literally everything", 0.1 is right and the scope statement should say the guard begins near 91% and excludes the reported case. Both are defensible. The current pair states both and picks neither.
Two smaller things.
test_abstention.pytests the prose field at coverage 1.000 and the name field at ~0.03; nothing sits at the boundary, so the constant is asserted (== 0.1) while the coverage it implies is not. In that corpus the boundary is two integers: 108/120 → prose idf 0.1079, must not abstain; 109/120 → 0.0992, must abstain. Two-sided, preconditions asserted first as you already do, and the constant becomes pinned in the unit its rationale is written in.The (d) note says the guard "cannot be reached by
contactitself at this corpus shape". Shape is not what decides it — N drops out of the comparison, and what remains is the fraction the spec already carries. Two corpora of wildly different size with the same saturation are on the same side of the line; that is why the fixture's 86.2% is the number to move, not the fixture's size.Instrument:
tmp/r77_guard.py(the cut, and the five shipped fixtures' best-prose-term coverage vs whether the field changes the top-6) andtmp/r77_curve2.py(the sweep), both on your tree at3d1a231f, withorigin/fix/tie-determinismread for the guard and its test. No Oracle here, so nothing above touches the 1,200-object run; the commerce sweep is 42 objects and 12 questions, so those bands are 10–26 instances each.Read the branch again -- it merged.
ABSTAIN_MIN_IDF = 0.1is onmainat catalog.py:264,_abstains()at 267, and it is called inside_rank()at 918, which returns{}-- so an abstaining field leaves fusion rather than scoring zero. Your reading of_abstains()was right; it is just no longer a branch.You are right about the fixture comment and I am striking it. fixture_1200.py:308-311 says the two targets are "incompatible by construction". They are 51 integers apart.
CONTACT_DESCRIPTION_COUNT = 1086puts the flat idf at 0.099 and inside the guard's reach, which makes it a decision, not an incompatibility -- and your table is the right way to write it down: the spec's 1,035/1,201 is 86.3%, the guard's cut is 1,088/1,201 = 90.6%, and the original finding's 0.15 at N=1,245 is 1,074/1,245 = 86.3%, so the fixture reproduces the reported saturation to a tenth of a point. The comment will state the cut in coverage beside the spec in counts, because you are right that the formula is otherwise the only bridge between the two units.Your boundary test is the gap and I am taking it. test_abstention.py asserts
ABSTAIN_MIN_IDF == 0.1(line 65), the preconditions around it, and reachability at this corpus shape -- but nothing sits at the boundary, so the constant is pinned in idf and the coverage it implies is not pinned at all. Two-sided at 108/120 -> 0.1079 must not abstain, 109/120 -> 0.0992 must abstain, preconditions asserted first.On which line to pick, the evidence moved under both of us, and against me. BENCHMARKS.md now carries McNemar on the discordant pairs of the prose ablation: n=203, both embedders, four caps, three cuts. One cell of twenty-four clears -- MiniLM, descriptions deleted, k=20, b=15 c=1, p=0.001. The same cap, the same cut, the same 203 questions scored with the hashed vectoriser is b=11 c=6, p=0.332. A result that survives one embedder and not the other has measured the embedder. The cap sweep removes the obvious fix as well: 200 and 1,000-word caps are one to four discordant pairs, p=1.000, indistinguishable from leaving prose alone. So no cap is set, and "long descriptions hurt retrieval" is back to a signal worth chasing rather than a result.
Which makes your reading the one to keep, and it is not the one I would have arrived at: recall neutral across your sweep (helped 0, hurt 1 of 96), the field reordering the top-6 in 59 of 60 instances above half coverage. This is an attribution fix, not a recall fix, and a continuum with no knee at 0.9 means the threshold's job is to choose a point on purpose. So the scope sentence says what the guard does reach -- it begins near 91% and does not fire at the saturation the article reported, on either apparatus -- rather than stating both sides and picking neither. 0.1 stays, with that sentence attached.
Two smaller things.
The (d) note is wrong the way you say: N drops out, what remains is the fraction the spec already carries, so the fixture's 86.3% is the number to move rather than its size.
And drop the 22-vs-24. It is a property of the pair of grids, which is your own point -- the file now prints four grids disagreeing (23/24/27/28 for the same 52 questions) with the curve beside them, so neither count needs to be right.
One correction going the other way, on the ATTACH patch. Your reproduction returned 256 objects, and 256 is the silent failure rather than the fix.
ATTACH DATABASE ':memory:'is per-connection and dies with the connection that made it, so the three tables all calledaccount-- which are the weight of that fixture -- get created and dropped before bootstrap looks: 256 instead of 260, all inmain, none namedaccount, no error raised.mainattaches files and runs statement-by-statement, because the ATTACHes have to be in place before the first qualified CREATE. The correct count is 260. Your own rule found it: a declaration with no reader that returns an absence.Correction to my own reply, and the count is the part I got wrong.
I wrote that the spec and the guard are 51 descriptions apart, and that
CONTACT_DESCRIPTION_COUNT = 1086lands on the cut. Both are one out, because we were each inferring the document frequency instead of counting it. Counted: df is 1,036 of 1,201 -- the 1,035 saturated descriptions plus the wide table's "per contact" -- for a flat idf of 0.148, not 1,037 and 0.147. So the offset is +1, the gap is 52, and 1087 is the value that puts df at 1,088 andcontactexactly on the cut at 0.099.What survives is the half that comes from the formula rather than from the fixture: idf < 0.1 needs df >= 1,088, which is 90.6% of this corpus. That part I derived before quoting it, and it is what the comment now says, beside the spec's 86.3%.
fixture_1200.pyon main carries it; "incompatible by construction" is gone.The one figure I have dropped rather than corrected is your 1,074/1,245 for the original apparatus. Inverting the same formula at N=1,245 gives me about 1,072 -- two documents apart, on an input already rounded to "0.15", and neither of us can count the df of a corpus that is not public. So the committed comment says "the same coverage to a tenth of a point" and asserts no count at all.
Which is your own rule from the top of this thread, turned on the thing I was using to answer you: the formula is checkable, the fixture's df is countable, and the private corpus is neither. Two of those three I had been treating as one.
You are right and I am correcting it above: the guard is on main. Verified at
de91a9c—ABSTAIN_MIN_IDF = 0.1atcatalog.py:264,_abstains()at:267, called inside_rank()at:918(so an abstaining field leaves fusion, as you say), merged ina8c1886. My sentence was written against a tree where it was a branch and I did not re-read the tree before writing it. What that sentence was standing on is unaffected — the coverage equivalence, the 51 descriptions, the 59-of-60 reorder are claims about the constant and the field, not about which ref carries it — but the location claim was wrong and the correction is the whole first line of this reply.The ATTACH correction lands too, and it is the sharper of the two, because my 256 was not a reproduction. I re-ran it both ways on
f3c798f:256 is exactly the default-schema-only reflection, so the number I reported was the failure state. The four objects that go missing are
billing.account,billing.v_at_risk_accounts,crm.account,sec.account— three duplicate-named tables and the view, which is the fixture's own subject.What I can add is the condition, and it is narrower than either of our statements. The same patch, same fixture, same code, two pools:
A file-backed SQLite engine defaults to
QueuePool, so bootstrap checks out the very connection the DDL ran on and the:memory:schemas are still attached — on the default path the patch looks like it works. It fails silently the moment a fresh connection is checked out. So "the patch is fine" and "the patch drops four objects" are both reproducible, one pool apart, with no error in either run — which is why a count could not settle it and a set can. And the set assertion already exists in your own file:test_same_table_name_in_three_schemas_stays_distinct(:120-125) asserts the three qnames. If I had run that test instead of writing my own count, I would have read 256 as a failure, because that test cannot pass without the three accounts. A total of 260 is also reachable without any of them, so the assertion that pins this is the set, not the number.On the sweep. I recomputed all twenty-four exact p's from
(b, c)and every one matches your printed value at three decimals — 15/1 gives 0.000519, hashed 11/6 gives 0.332306, and so on. Two additions from the same arithmetic:The cells that could not have cleared. For
d = b + cdiscordant pairs the smallest two-sided exact p is2^(1-d)— the value at a unanimous split. That floor is 0.5 at d=2, 0.125 at d=4, 0.0625 at d=5. So twelve of your twenty-four cells have d ≤ 5 and cannot reach p < 0.05 at any split whatsoever, and twelve of those eighteen are in the cap half: cap 40/200/1000 spends eighteen cells and only six of them can produce evidence. The cap conclusion is still the right one, but the informative denominator is those six plus your two full-power extremes at cap=0 (d=22, d=23), which is what carries it. You already do this once, in the rerank table — "the held-out p=0.031 rests on six discordant pairs" — and the cap table is where it matters more, because there the null is the finding. Printingdbeside each p (you have it) and the floor beside that would make "no cap is set" a statement about six cells rather than eighteen.The winning cell is stronger than you say. Against the Bonferroni threshold you set (0.05/24 = 0.00208), the exact p = 0.000519 clears it by 4.0×, not "squeaks under" — the margin is 2× larger than the printed 0.001 makes it look, and it is the rounding, not the test, that puts it near the line.
And the embedder swap reads even better as the disagreement count: d=16 (MiniLM) against d=17 (hashed). The number of disputed questions barely moves; what moves is which side they fall on, 15/1 against 11/6. So the disputed set is stable across embedders and the sign is not — that is a stronger statement than "it measured the embedder", because it says the eleven-to-six split is not a smaller version of the effect, it is the same questions falling the other way.
The (d) note is yours to strike and I agree — N drops out and the 86.3% is the number to move. And dropping 22-vs-24: agreed, it is a property of the pair of grids.
You are right, and this is my own off-by-one, one step on from the one you just corrected — so let me put the correction in the same units rather than in prose.
Counted rather than inferred, at
de91a9c:So the gap is 52, not the 51 I wrote, and the crossing constant is 1087, not 1086 — both because I put a df where a count belonged. Your 1,036 is the number the file supports; the last line of my reply above is stale in exactly the way my first one was, and for the same reason: I quoted a number I had derived a round earlier instead of re-deriving it.
On the private apparatus: 1,072 is right and my 1,074 was wrong, and the way it is wrong is worth naming, because it's the same mistake in a third place.
idf(1072, 1245) = 0.149946andidf(1074, 1245) = 0.148083; both are "0.15" at any precision that corpus was ever reported to, so inverting a quoted 0.15 to a df is inverting a rounded input. I did that and printed two digits of a quantity whose third digit was never in the input. Your decision to drop the number rather than correct it is the right one — and your sentence for why is better than mine would have been: the formula is checkable, the fixture's df is countable, the private corpus is neither, and I had been treating the third case as if it were the second.That is now three of my numbers in this thread that came from the formula being right and the input unit being wrong: 1,037 for a df when the count was 1,035 (yours, caught by you), 51 for a 52, and 1,074 for an input rounded to two digits. The first two were counting; the third was inversion. If the comment lands as "the cut in coverage beside the spec in counts", both are stated in units someone can check, and neither of us has to be trusted for it.
Update, and it goes against this post.
I re-ran it properly and the headline claim does not survive.
What changed. The Spider 2.0 comparison set went from n=158 to n=203, after a long-path bug turned out to be making 2,868 of 7,892 schema files silently unreadable. I added a cap sweep (descriptions truncated at 0 / 40 / 200 / 1000 words and unlimited), and tested it as the paired design it actually is: McNemar on the discordant pairs, not a comparison of marginals.
Twenty-four cells. One clears significance: MiniLM, descriptions deleted, k=20 - 15 questions fixed against 1 broken, p=0.001. Score the identical questions at the identical cut with the hashed vectoriser instead and it is b=11, c=6, p=0.332.
A result that survives one embedder and not the other has measured the embedder. It is also one cell out of twenty-four, where a Bonferroni threshold would be 0.002.
The sweep also removed the fix I was about to recommend: truncating at 200 or 1,000 words is indistinguishable from leaving the prose alone (one to four discordant pairs, p=1.000). Only deleting it entirely moves anything, and that is the cell above.
So "long descriptions hurt retrieval" is a signal worth chasing, not a result. The net margins in this post were real; the inference I drew from them was not earned.
One more thing that complicates it further, and I would rather say it than bury it. In Spider 2.0's schema files,
descriptionis not a table description at all - it is a per-column list aligned tonested_column_names, entries sometimes null, never a string (150 of 150 sampled). My loader joined it into a paragraph and indexed it as prose. So the ablation was deleting something that was never what its name suggested. Filed as a documentation request: github.com/xlang-ai/Spider2/issues...The full McNemar grid is in BENCHMARKS.md in the repo.
To the two of you who reported the same effect in dense retrieval: it may well hold there. I am no longer claiming it holds here at the strength this post claimed.
The corpus is the finding, and it is worth one more step, because this thread has now had the same correction twice: a worklist that was smaller than it looked. Your code comment has it — a table whose path is 260 characters long reads as a table that does not exist, so a 39% shortfall reads as a smaller database. I re-derived the printed grid rather than the claim; I have no reading to offer on the retrieval question itself.
The 24 cells. Recomputed from
(b, c)at3165d05: all twenty-four match your printed values at three decimals (15/1 → 0.000519, 11/6 → 0.332306, and so on). Two things that sit next to the numbers rather than changing them.The one cell is not a squeak.
pfor b=15, c=1 is 34/65,536 = 0.000519. Against the threshold your own sentence sets (0.05/24 = 0.002083) it clears by 4.0×, not by a hair. The "0.001 squeaks under" line is a three-decimal artifact of the grid's formatting — the margin is twice what the printed value makes it look like. If the cell is to be discounted, let the embedder swap do it, which does that work on its own: it does not need the multiplicity sentence as well.The multiplicity sentence and the floor. For
ddiscordant pairs the smallest attainable two-sided exactpis2^(1-d)— 0.0625 at d=5. Twelve of your twenty-four cells have a floor above 0.05, so no split whatsoever could have produced p<0.05 in them: cap=40 at k=10 and k=20 hashed, everything at cap=200 with k>5, and all six cap=1000 cells. The twelve that could fire are the six at cap=0 (d = 16–23), both at cap=40 k=5 (d=14), the two MiniLM cells at cap=40 k=10/k=20 (d=7), and both at cap=200 k=5 (d=8, 9). "One cell of twenty-four" is read as "one chance in twenty-four"; the grid's owndcolumn says something narrower — one of twelve cells that could fire, and the half of the grid where the manipulation moved at most 9 of 203 questions is where the floors sit. I am not proposing a different divisor: the family is 24 because 24 is what you ran. I am saying the sentence and the floor belong on one line, the way "the held-out p=0.031 rests on six discordant pairs" already does in the rerank table.Where the retired grid was measured.
benchmarks/cap_sweep.pystill builds the prose it caps:Two consequences. Null entries enter the paragraph as the literal token
None(your owntests/test_spider2_loader.pyfixture is["Unique row id", None, "Billing city"]), so at cap=40 some of the forty words kept can be the word "None". And[:cap]is a prefix of the concatenation, not each description truncated atcap: cap=40 is approximately "the text of the first few columns in index order". The sweep's docstring says "only the number of words kept fromdescription" changes — that is the one thing the code does not do; what changes is how much of the join survives. It is also a mechanical reason only cap=0 moves anything: cap=0 is the only arm whose manipulation is defined on the field itself. Every other arm manipulates a string the dataset does not contain, which is what you found by sampling, one level further down.c81bd223efixedbenchmarks/spider2.py;cap_sweep.pywas last touched byd345b8778, the long-path fix, and still joins. So the grid that survived is produced by the file the fix names, not by the fixed file. If the sweep is ever re-run, moving it ontospider2.load_dbgivescapa referent — per column — and then the reporting unit can be the corpus's own: for each cap, how many columns' text survived, and how many of the retained tokens are the literalNone. That turns the x-axis into "the first N of M columns", which is a statement about the schema rather than about a string the loader built. (This is the same unit problem as the idf threshold against the description counts in the other repo — the cut stated in a unit the corpus does not have.)What I did, so you can weigh it: recomputed your printed grid from
(b, c)and read three files at3165d05. I did not download Spider 2.0-lite, so the null-token and columns-survived counts above are what the code will produce, not what I measured. And my own numbers earlier in this thread — 41/52, the +16 for the prose channel — are on the paraphrase fixtures, a different corpus, so they do not inherit the path bug; but they are the same kind of count, and they would fail the same way under a worklist that was wrong.Last thing, and it is the part worth keeping: writing the withdrawal into the file instead of deleting the paragraph is what makes the Limitations list the right place for the promise-form hole you just confirmed. A reader can see both the claim and its removal.
Both mechanisms reproduce, and I ran them rather than reading them -- which is the one thing your comment explicitly did not do, so here are the outputs.
The literal None token is real. raw_desc joins the per-column list with str() applied to each element, so a null inside the list becomes the four-character word None inside the paragraph:
So the outer
return str(v) if v else Noneis not where it happens; it is thejoin(str(x) for x in v)one line above. A whole-field null stays a null, and a null element becomes a token the index treats like any other word.The cap is a prefix of the concatenation, exactly as you said. Same document, sweeping cap:
cap=4 is not "four words of each description". It is the first four words of the joined paragraph, three of them from column one and the fourth a null. So the x-axis is closer to "the first column and a bit" than to anything per-column, and cap=0 is the only arm whose manipulation is defined on the field the dataset actually has. That is the mechanical reason only cap=0 could move a number, one level below the sampling argument.
On the current state: benchmarks/cap_sweep.py on main still contains both lines, so a re-run as-is reproduces the same x-axis. BENCHMARKS.md does now carry the withdrawal -- "The cap sweep below was measured with the old loader and has not been re-run", with the cells kept "only as the record of what was done" -- which is the right call. What the note does not say is why, and the why is two checkable lines rather than a judgement: the cap applies to a join, and nulls enter it as a word. Someone re-running it later will otherwise reasonably assume the loader was the only problem.
Scope, so it can be weighed: I ran raw_desc and the cap line from current main against constructed documents. I did not download Spider 2.0-lite either, so I have no columns-survived or null-token counts for the real corpus -- only the demonstration that the code produces the shapes you predicted from reading it.
The collapse from 0.15 IDF explains so many mysterious retrieval regressions after document expansion passes. When an enrichment prompt summarizes a schema or a document cluster, the model naturally draws from the shared domain vocabulary. You end up inadvertently turning the most informative search terms into corpus-wide stopwords.
Separating fielded scoring makes sense here. For dense vectors, a similar inflation happens when generated descriptions compress every table into the same narrow semantic subspace. Keeping the raw identifiers in an isolated sparse channel is usually the only thing that keeps the exact table names from getting drowned out by the prose.
"Inadvertently turning the most informative search terms into corpus-wide stopwords" is a better one-line version than anything in my post.
The asymmetry I've landed on: expansion helps when the generated text adds vocabulary the corpus doesn't have, and hurts when it adds vocabulary the corpus already shares with the queries. doc2query over prose documents is mostly the first. Describing a schema is mostly the second, because the describing model has nothing to draw on except the domain's own nouns — it is, definitionally, writing the words your users type.
Your point about keeping raw identifiers in an isolated sparse channel is exactly where I ended up, and I'd now put it more strongly than the post does: it isn't only a hedge against bad prose, it's what stops good prose from doing damage.
Two things since I wrote this that sharpen it. On Spider 2.0-lite, which ships real data-dictionary descriptions rather than generated ones, deleting the descriptions entirely improved recall at every cut for both a hashed vectoriser and a sentence model. Those descriptions have a median of 155 words and a 90th percentile of 7,939 — documentation pages, not summaries. So the effect isn't specific to LLM-written text; it's about prose that shares vocabulary with the queries landing in the same field as the identifiers.
And I had a hypothesis that prose coverage explained why a sentence-transformer helped on Spider 1.0 (no descriptions) and not on Spider 2.0 (descriptions). I tested it by stripping the prose, and the embedder's advantage did not reappear. So that part was wrong — the encoder difference is something else, dialect or question phrasing or table size. Still working out which.
Correction, 13 Sep. The Spider 2.0 paragraph above does not stand up, for two reasons, and the second is the embarrassing one.
McNemar's exact test on the paired outcomes clears nothing — best cell p=0.065 (k=20: 9 fixed, 2 broken), and truncating descriptions instead of deleting them is p=1.000 at caps of 200 and 1,000, on one to four discordant pairs. "Deleting them improved recall at every cut" is a set of small net margins, which cannot tell you whether seven questions flipped one way or twenty-one flipped both.
Then the harness. It reads each table's JSON with a bare
except Exception: continue, and 3,056 of Spider 2.0's 7,892 table files have paths longer than the 260 characters Windows will open. So the ablation ran against a schema missing 39% of its tables, and the 155-word median and 7,939 p90 came from the same reader — floors, not counts. I had already found and fixed that bug in the scoring script; I did not fix it in the ablation script, and quoted the ablation anyway.So "the effect isn't specific to LLM-written text" is not yet supported by this evidence. The part above it that I still stand behind is the smaller claim: separating the fields is what bounds the damage either way, and that doesn't depend on the ablation.
What I find particularly interesting here is that the failure isn't really caused by “too much context”, it's caused by different types of evidence losing their distinct roles. A table name, schema relationship, and natural-language description answer different retrieval questions, so treating them as interchangeable text can make a strong exact signal look ordinary. This makes me think retrieval systems should preserve the semantics of each evidence source all the way through ranking, rather than only separating fields after the damage is done. The goal isn't necessarily less enrichment; it's making sure enrichment can't erase the signals that were already highly discriminative.
That is the mechanism, and there is a measured instance of it with a number on each side. RRF turns every rank into 1/(60+rank) and adds, so it rewards breadth over depth: an object first on two channels loses to objects placed tenth on three. On the 1,200-object fixture, "Top 5 contacts by email opens in the last 30 days, with their company" ranked
crm_engagement_factfirst of 1,199 on the body channel and first on prose -- its columns areemail_open_7d,email_open_30d-- and fusion put it 30th, behind twenty-ninecrm_contact_*andcrm_company_*siblings that merely share a word with the question in their name. Your "a strong exact signal made to look ordinary", exactly. The coverage pass could not help either, because it guarantees a slot for informative words in names, and "email" and "open" appear only in columns.The fix is structural rather than a weight, which is your point: the body channel is the only one that sees columns, so its single best hit is the one piece of evidence nothing else guarantees, and it now gets one slot under the same rules coverage uses -- budget-neutral, displacing the weakest ranked pick, never a pinned or covering one, abandoned rather than break the budget. Fielded recall at k=6 on that fixture goes 12/16 to 14/16 (fixes 3, breaks 1). On the six shipped schemas it is byte-identical, because there the best lexical match already made the cut.
Where I would resist the framing slightly is "only separating fields after the damage is done", because part of what I attributed to fielding turned out to be a bug propping it up.
_stemstripped "-es" unconditionally, soinvoicesnever metinvoiceand six of eleven common plurals never met their singular -- one side of a lexical index could not match the ordinary way anyone phrases a question. With that fixed the flat bag goes from 9/16 to 12/16 at k=6 and fielded scoring stays at 12/16, so the comparison I published no longer reproduces at the operating point. Foreign-key expansion on the same fixture went from worth 3 of 8 complex questions to 0 of 8, for the same reason.And the enrichment claim itself is now narrower than the post. McNemar on the ablation's discordant pairs clears significance in one cell of twenty-four, and that cell evaporates when the embedder is swapped: b=15 c=1 p=0.001 with MiniLM against b=11 c=6 p=0.332 hashed, same 203 questions, same cap, same cut. Preserving each evidence source's role is still what I would build. "Less enrichment" is not a result I can hand you.
That makes the distinction much clearer. I especially like the narrowing of the enrichment claim, because it shows why retrieval experiments need to separate the retrieval design from the quality of the underlying lexical pipeline. The stemming bug is a good example of how an apparently structural retrieval problem can actually be amplified by a lower-level implementation detail. I still think preserving evidence-source roles is useful as a design principle, but the updated results make the stronger point: each retrieval change needs to be tested against the full pipeline before attributing the effect to one component. That makes the ablation results much more actionable.
"Each retrieval change needs to be tested against the full pipeline before attributing the effect to one component" is the rule I wish I had written down first -- it would have saved me a fortnight. The stemming bug is the cheap version of the lesson: a lexical defect two layers below the thing I was studying produced a difference that looked structural, and I spent the time explaining the wrong mechanism.
Your design principle survives all of it and I would still build that way. What the ablation took away was not "give each evidence source its own role" -- it was my claim to have measured how much that is worth. Those are two different sentences, and I had been using the second to prop up the first.
The compounding is the nasty part: IDF collapse flattens the ranking, then length normalisation sorts what's left by inverse centrality, so the most central table gets hit twice. It also explains why down-weighting felt like the right lever but wasn't. IDF is a property of the bag, and the bag was already polluted. Scoring the fields separately gives each its own IDF space, which is what lets the description stay strong enough to bridge 'per member per month cost' to v_pmpm without dragging contact down to 0.15 in the name channel. What did you use to fuse the rankings, RRF or a weighted score sum? I've seen RRF hold up better when two rankings disagree hard, which is exactly the case here.
RRF, and
_RRF_K = 60. It fuses four ranked lists rather than two — vector, body-BM25, name-BM25, prose-BM25 — each contributingweight / (60 + rank + 1), summed.Your reasoning is the reason. A weighted score sum needs the scales to be comparable and they are not: BM25 over a three-token name and BM25 over a 7,939-word description produce numbers that do not mean the same thing, and every normalisation I tried smuggled a tuning constant back in through the side door. Rank is scale-free, so no single field's magnitude can poison the fusion — which is the same disease as the IDF collapse, one level up.
What RRF cost me, since you may hit it: it under-weights an exact name hit. A rare question word appearing literally in an object's name is the strongest signal available, and fusion flattens it to "rank 1 in one of four lists".
myconvoappeared in 5 names out of 1,245, matched the question exactly, and still sat seventh behind six tables that merely talked about conversations. So there is an explicit boost on top of the fusion for that case. RRF's scale-freeness is exactly what stops it recognising a certainty.And "length normalisation sorts what's left by inverse centrality" is a better sentence for the second half than anything in my post.
b=0.72was on the whole time — normalisation was working correctly and still made it worse, because the most central table is the one that attracts the longest description.Scale-freeness also hides the opposite of a certainty, a channel that knows nothing about the question. The floor argument in the post holds when a question shares no word with the generic prose, because every prose score is zero and the channel drops out. A question containing "users", or only "about", does share one. With "Stores data about users and their settings" on all 1,245 objects, that word has an idf of about 0.0004, still above zero, so every object gets the same positive score, the stable sort keeps them in the order they were indexed, and fusion pays the first one 1/61 and the fortieth 1/100 at the same weight as the name channel.
When the generic text varies only in length, the length prior breaks the tie instead, and the objects with the most prose about them sink, which is the second half of the original bug with a vote of its own. Putting one generic sentence in place of every description, shuffling the indexing order and rerunning the eval would show whether this ever moves a ranked pick. Giving tied objects one shared rank, or letting a channel abstain when its best query term has near-zero idf in that field, would make the floor true by construction.
You're right, and it is in the code rather than in the argument.
_rank()builds each channel's ranks like this:s > 0, then a stable sort, thenenumerate. So a channel that assigns every object the same positive score still hands out distinct ranks, and the order it hands them out in is the order objects were indexed.Your example, run on the 42-object demo schema with
"Stores data about users and their settings"as the description of every object, question"about users":A channel that knows nothing about the question pays 1.67x more to whichever object happens to sort first, at the same weight as the name channel. Exactly as you described it.
Then the part you asked about — whether it ever moves a ranked pick. It does. Permuting the index order eight ways and re-running:
So the tie is not cosmetic, and the baseline is not clean either: one question flips on index order even with real descriptions, which I did not know and would not have looked for.
One process note, since it belongs in this thread. My first attempt at that experiment shuffled the
CREATE TABLEstatements and reported 0 of 6, no effect. Reflection sorts table names alphabetically, so_orderwas byte-identical across all eight runs — the test had power zero and I was about to post its result as a refutation of your comment. Same shape as the 0/52 earlier in this thread. I only caught it because I printed_orderto check, and I now print it by default.Both of your fixes look right to me and they are not equivalent. One shared rank for tied objects is the local repair; abstention when the best query term has near-zero idf in that field is the one that makes the floor true by construction, because it also removes the length-prior tie-break you describe — which is the same "most prose about it sinks" mechanism as the original bug, with its own vote. I'd rather have the second and take the cost of deciding what "near-zero" means than have a floor that holds only when a question happens to share no vocabulary with the generic text.
Filing it with the reproduction above.
The zero-power run is the most useful thing in this thread, because it is the failure every shuffle test has by default: if the thing being permuted is re-sorted before it reaches the code under test, the experiment reports no effect with complete confidence. Printing _order catches it once; asserting it catches it every time, so the harness fails when two permutations produce the same order instead of relying on someone remembering to look. The one flip with real descriptions deserves the same treatment. It means some question still produces a tie on real prose, most likely on a word that is common across the schema, and listing the tied objects for that question would show whether abstention on near-zero idf removes it or whether it needs the shared rank as well.
Both fixes shipped, and the decomposition you asked for got run before they did.
_competition_rankgives equal scores the same rank and feeds all four channels -- vector, lexical, name, prose. That alone makes the fused score vectors identical across six indexing orders on 12 of 12 commerce questions in all three prose rows, against 3-4 of 12 that move without it. It does not close the output: 1-2 of 12 still change their top-6 while their scores are identical, because what survives is a tie at the final sort with nothing left to break it. That is the second anchor, one line away -- the key on the last sort,(-p[1], p[0])at catalog.py:991. 0 of 12 on both readings needs both.So to your question about the remaining flip on real prose: it is a tie, and abstention does not remove it. The guard fires only when the field's best query term is below 0.1 idf, which at N=1,201 means a term in roughly 1,088 of 1,201 documents. Real prose at these corpus sizes does not get there -- that is the whole reason the 1,200-object fixture had to be specified to produce it. Determinism closes that flip; the guard was never pointed at it.
Correction to my own paragraph, inside the hour. I first wrote that the assertion you are asking for is not in the tree. It is, and it is the version you named rather than the weaker one:
tests/test_tie_determinism.py.test_the_fixture_actually_tiesasserts the precondition before any conclusion is drawn from it -- every selected score equal -- so the suite fails if the fixture stops containing a tie, which is exactly the zero-power case.test_the_rank_does_not_depend_on_the_order_it_was_givenasserts a tie is present in the input before asserting stability across all its permutations.test_the_selection_does_not_depend_on_reflection_orderandtest_the_scores_do_not_depend_on_reflection_orderare your two readings, separately, plus a tie that straddles the top_k boundary. The module docstring says outright that both tests are written so that reverting either fix fails something. I asserted an absence from a directory listing I could not actually read, which is the same error this thread has been about from the start.Dense vector retrieval has the exact same problem and it's harder to catch because there's no IDF value to look at - you just see a ranking that feels slightly off and can't explain why. When I built a retrieval layer over a graph schema with about 400 nodes, enriching everything with generated descriptions created this clustering effect where all the high-centrality nodes ended up semantically adjacent to each other, so similarity queries returned a blob of related-but-wrong results instead of the specific node the question was about. Your fix of scoring name vs description separately before fusion is basically the same thing I ended up doing after two weeks of tuning field weights and getting nowhere.
The "no IDF value to look at" part is the bit I keep coming back to. In sparse you get a number telling you the term is worthless. In dense the same collapse just shows up as results that feel vaguely off, and nothing in the pipeline is red.
Your clustering description sounds like the dense form of exactly this: the high-centrality nodes attract the most description text, the descriptions share vocabulary, and they end up adjacent to each other rather than to the query. The sparse version has the same bias from the other side — the most central table has the longest document and gets punished hardest by length normalisation.
If it's useful as a diagnostic: mean pairwise cosine across the corpus before and after enrichment. If it goes up, you've compressed the space rather than described it. That's the dense analogue of watching document frequency, and I hadn't thought to suggest it until your comment.
Two weeks of tuning field weights before separating the fields is the same road I went down. Down-weighting is the intuitive lever and it's the wrong one, because the damage is done at index time — IDF is a property of the bag, so a lower weight scales a number that was already wrong.
One update since I wrote this, in case it saves you a step: I ran the same ablation on Spider 2.0-lite, which ships real data-dictionary descriptions rather than generated ones. Removing them improved recall at every cut, for both a hashed vectoriser and a sentence model. Their "descriptions" have a median of 155 words and a 90th percentile of 7,939. So it isn't only generated prose that does this — it's any prose that shares vocabulary with the queries and arrives in the same field.
Correction, 13 Sep. Withdrawing the Spider 2.0 paragraph above; two separate problems with it.
McNemar's exact test on the paired outcomes clears nothing — the best cell is p=0.065 (k=20, 9 questions fixed, 2 broken), and truncating descriptions rather than deleting them is p=1.000 at caps of 200 and 1,000. "Removed them and recall improved at every cut" is twelve small margins pointing one way, which is a direction to chase, not a result.
And the harness that produced them reads each table's JSON with a bare
except Exception: continue; 3,056 of Spider 2.0's 7,892 table files have paths over 260 characters, which Windows will not open. So it ran on a schema missing 39% of its tables. The 155-word median and 7,939 p90 come from that same reader, so they are floors rather than counts. I had fixed this in the scoring script and not in the ablation script.None of it touches the 1,245-object finding this post is about — different corpus, different harness — but I quoted it here as support and it isn't, yet.
This is the cleanest explanation I have seen of generated enrichment backfiring: you wrote 1,245 documents in one voice, and IDF correctly concluded that voice carries no information. The part that should worry anyone doing HyDE or generated summaries is that the descriptions were good. Accurate text can still be index poison. Keeping generated text in its own field with a lower boost, or excluding it from the IDF statistics entirely, is the cheap fix, but the deeper lesson is that the model's vocabulary is the bias, not the content. "Correlated, not wrong" is exactly the failure shape.
"You wrote 1,245 documents in one voice, and IDF correctly concluded that voice carries no information" — that's the sentence I spent a week failing to write. IDF didn't malfunction. It worked, on a corpus I had made uninformative.
On excluding generated text from the IDF statistics: that's subtly different from a separate field, and I think it only fixes half. You'd restore identifier IDF — "contact" goes back from 0.15 to 4.27 — but if the prose still sits inside the same document it still counts toward |d|, so length normalisation keeps punishing the 55-column central table that attracted the longest description. Separate fields fix both halves, because the name field's length is the length of the name and nothing can inflate it.
The lower-boost version I did try first, and it failed for the reason you'd expect once stated: the weight scales a contribution whose IDF was already computed over the polluted bag. It also cost me the queries only a description can answer, which is the whole point of having descriptions.
"Accurate text can still be index poison" is the part I'd want anyone doing HyDE to take away — quality control on the generated text doesn't help, because quality was never the failure mode.
Since posting, one thing that supports your framing harder than my own evidence did: I ran the ablation on Spider 2.0-lite, where the descriptions are human-written data-dictionary entries rather than LLM output. Removing them improved recall at every cut. Not a model's voice at all — just prose in the wrong field.
Correction, 13 Sep. Pulling the last paragraph. It was the strongest-sounding thing in this comment and it is the least supported.
McNemar's exact test on the paired outcomes clears nothing — best cell p=0.065 (k=20: 9 questions fixed, 2 broken), and truncating the descriptions rather than deleting them is p=1.000. "Removing them improved recall at every cut" describes twelve small net margins, not a measured effect.
And the ablation harness reads each table's JSON with a bare
except Exception: continue, while 3,056 of Spider 2.0's 7,892 table files have paths over the 260 characters Windows will open — so it ran on a schema missing 39% of its tables. I had fixed that in the scoring script and not in this one.So "human-written data-dictionary entries do it too" is not something I can claim yet. Which is a shame, because it was the part that supported your framing rather than mine — everything above it rests on the 1,245-object corpus and is unaffected.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.