Hey everyone. This time I'll go through what got me started on this benchmark, and the core of how it's actually built.
It started from reading co...
For further actions, you may consider blocking this person and/or reporting abuse
You already have the number that answers your own question, in the record you linked. I downloaded the eight LongMemEval runs from
beam1m-tablet-2, reproduced all eight scores exactly (95.2 / 96.0 / 96.0 and 93.2 / 94.0 / 93.8, with theok: nullconvention leaving one question out of the three GPT denominators), and then computed the thing the paper can't compute from the aggregate: the per-question pairing.Reader swap vs. a plain re-run, per question. Over the nine opus↔gpt run pairs, the two readers disagree on about 27 of 500 questions (4.8–6.4%). Opus wins 171 of those and loses 75, so the direction is solid: +2.14pp, se 0.35pp, z = 6.1. But now the same statistic for two runs of the same reader: 12–26 questions flip, about 19 of 500 (2.4–5.2%). So the reader effect is roughly 1.4× the run-to-run noise, and a single pair of runs is what a leaderboard row actually is.
That gives the number a version is missing — its resolution, published next to the reader and the judge:
The per-pair standard deviation of a gap is 1.05pp, so anything closer than ~2pp at one run per system is inside the noise. Your own two LongMemEval rows for tablet-2 differ by exactly 2.0pp (95.7 vs 93.7) — the row you use to make the reader visible sits on the resolution line. Compute this once per version (re-grade a fixed slice with a second reader, pair it per question) and "five is certified" starts to mean something specific: certified above the resolution, not merely five runs.
And the band is not the reader's alone, which changes how it can be quoted. Conditioning on your own
gold_in_poolfield, on the ~376 questions where the engine delivered the evidence, the reader swap is worth +0.8 to +1.9pp (3.5–4.8% of questions flip). On the ~123 where it did not, it is +3.3 to +6.5pp (6.5–11.3% flip). The reader effect lives almost entirely in the questions where retrieval failed and the reader fell back on its own priors. So it is an interaction — reader × your engine's failure rate — and it will shrink for a better engine and grow for a worse one. "That gap belongs to the reader" is the right arc in the claims table, but a swap number measured on tablet-2 will not transport to a competitor's submission. Two ways out, both cheap: publish the band per submission, or publish the gold_in_pool split, which you already record.Three things I'd want before believing a result, since you asked:
glasshouse-v0.1/README.md— the reader and judge are named, but how much either of them moves a fixed set of answers is not.headline.scoreandheadline.stdev— aggregate against aggregate, which throws the pairing away and is roughly twice as wide as it needs to be. Comparing two systems on the identical question set, counting the discordant questions and stating the gap in questions rather than points is a two-line addition to whatever reports the headline, and it is the difference between "1.2pp apart" and "8 questions apart".The limitation on my numbers, stated rather than left implied: the two reader groups are different runs, so my pairing controls item difficulty but cannot separate the reader from the run draw — that separation would need the same answers graded twice. That is not a criticism of the record, it is why I compared the swap against the same-reader re-runs, which is the piece of your design that made it measurable at all.
To your second question: the tip-off is a table where one system has three numbers and none of them carries a resolution — the LoCoMo row in your own README (84 / 58.44 / 75.14) is that table. Not dishonesty, just a benchmark whose noise floor was never measured, so every reading had to be argued instead of settled.
Want this as a PR under
individuals/? Your own README says a recomputation is evidence rather than a proposal, and it should go to the submission it disputes — so tell me which of the two you'd prefer and I'll put it where it belongs. The script that produced the table is a couple of dozen lines and I'm happy to hand it over either way.Thanks, this is really useful. Glasshouse v1 isn't in the repo yet, so there's nowhere for a PR to land right now. It's going up soon. I'll message you when it's there, and if you send the PR then I'll review it and update.
Good — and the wait is fine, I would rather send something checkable than something fast. When v1 lands I'll bring the run I already did against your published
beam1m-tablet-2records (the eight LongMemEval files): reproduce all eight headline scores, then the per-question pairing — reader-swap flips against same-reader re-run flips, and the smallest gap each run count can settle. That was the substance of my comment, so it should go where the number it disputes lives.One piece of it does not need v1, because it is a property of the version rather than of any run.
glasshouse-v0.1/README.mdnames the reader and the judge; what is missing beside them is how much either one moves a fixed set of answers. That is computable once per version — re-grade one fixed slice with a second reader and pair it per question — and it is the kind of number that is cheapest to write down before any submission exists, since every later row inherits it. Rule 8 says five runs is certified; a resolution figure is what would make "certified" mean something specific, and it would let a reader tell a 2pp gap at one run from a 2pp gap at five.If a spec sentence is easier than a per-version computation, the equivalent is to require submitters to report their discordant counts against a fixed reference run: you get the same number out of the submissions, and it costs you nothing to host.
Not expecting either in v1 — the version is drafted by companies and rule 10 sends additions to the next vote — just flagging what I would put in the PR so you know what is coming.
It's up. github.com/wontopos/glasshouse, under glasshouse-v0.1. Said I'd let you know.
The complaints list you started from is the real spec: published vendor numbers don't match measured ones, and swapping the grading model moves results more than the gap between the systems being compared - which makes most published numbers decorative. Writing the limits in the repo instead of promising them is the only credible move when you compete in the thing you administer. "We do not merge our own" is the sentence that matters, because it converts trust-me into check-me. The grader-model sensitivity is the sleeper issue in the whole space: if the judge choice dominates the system choice, the benchmark measures the judge.
You compressed my whole fairness section into one line, and the line is better
than the section. On grader sensitivity I should be straight: I don't measure it. There's a check that forces the string scorer and the LLM judge to agree (it caught an axis
where "I don't know" scored 0.5 one way and 0.0 the other), but that's not the same as swapping the judge and seeing how much moves. All I have is the per-question record, so someone can re-score with their own judge and publish the gap. Preferably someone who isn't me.
v0.1 is out: github.com/wontopos/glasshouse
The three-score stale fact axis is the bit I'd steal. Every binary eval I've touched mixes "no idea" with "confidently out of date" and you just can't see what's actually breaking in prod. The geography correction on latency: we ran a comparison this year where the Seoul-to-US round trip was 190ms and the gap between the systems was 40ms, so the whole result was just geography. Worth adding at some point: distributed evidence, where the right answer isn't in one turn but spread across five, because that's quietly where a lot of memory systems fall apart.
40ms against a 190ms trip is a better version of my own point. Thanks for that. You're right about distributed evidence. I counted: 3 questions out of 1,547 have the answer spread across five turns, and the gather axis is all two-turn. So it's basically not in there. I'll look at adding it myself. The benchmark isn't in the repo yet though, so once I've put the draft up, a PR from you on this would be appreciated.
v0.1 is out: github.com/wontopos/glasshouse
Publishing per-question records helps reproducibility, but the conflict of interest moves upstream: who decides which memory failures enter the 14 axes and how much each axis counts?
A benchmark can be reproducible and still favor the maintainer’s product strengths. I would want versioned rationale for every axis and weight, blind proposals from competing vendors, and a holdout set controlled outside the host.
Which part of the benchmark definition can a co-administrator actually veto?
Fair, and the weighting part I can answer now. Every axis has a written rationale behind its weight, including the one criterion I used to sort them and the two weighting schemes I tried and dropped. It isn't in the repo yet because what's up there is still a partial draft. When the full draft goes up that document goes with it, so as of right now you're correct that there is nothing for you to check. Blind proposals and an externally held holdout are both good and I hadn't considered either one. I'll look at both properly. On the veto question, you're right that there's no answer in the repo. That needs defining rather than hand waving, and I'd rather write it down carefully than give you something off the cuff here. Thanks for all three. If any of it ends up in the benchmark I'll come back to this thread and say so.
That sounds like the right sequence. Publishing the rationales and the weighting schemes you rejected makes the trade-offs inspectable; blind proposals and an externally held holdout reduce the ability to tune the benchmark toward the product. The remaining piece I’d put in the same governance note is change authority: who may alter axes or weights, what evidence is required, and who can veto a change after seeing results. Looking forward to the full draft.
The fourteen-day rule is the one I would look at again, because as written it hands you the advantage it is there to remove. Everyone gets a new version at the same moment, but a competitor has every reason to publish early, and the rule guarantees you fourteen days of iterating against a version whose leaderboard is already carrying other people's locked-in numbers. A submitter who publishes on day two cannot see what you are aiming at. You can see exactly what you need to clear.
The fix lives in the same family as your other rules rather than needing new machinery: hold every number on a new version for the same fourteen days and release them together, so nobody gets to tune against a published bar. The asymmetry right now is not that you publish late, it is that you are the only party who can choose to.
Good catch, and you're right that as written it works against itself. A submitter who publishes on day two locks their number in, and I get two weeks of looking at it before I have to commit to mine. That's the opposite of what the rule was for. I'll look at the rule again. If it changes, I'll come back to this thread and say what it became. Thanks for reading it closely enough to find that.
v0.1 is out, in case you want a look: github.com/wontopos/glasshouse
Building a benchmark for your own product is always a conflict of interest. The best approach I have seen is using a third-party eval suite as an independent sanity check.
v0.1 is out, and the weighting document went up with it: github.com/wontopos/glasshouse