Hey everyone. This time I'll go through what got me started on this benchmark, and the core of how it's actually built.
It started from reading complaints, not from an idea. The same ones kept coming up: numbers a vendor publishes don't match numbers someone else measures, swapping the model that does the grading moves the results more than the gap between the systems being compared, and because of that nobody really uses published numbers anyway. They test two options on their own data and keep whichever annoys them less. I work at Wontopos, and we sell a memory API, so I can't point at that and shrug.
All numbers below are from the current build. Nothing is final yet, so some of them will have moved by the time you read this.
Fairness first, because you have no reason to trust me
I know how "I built a fair benchmark" sounds coming from someone at a memory company. So instead of promising anything, here's what's written in the repo.
Wontopos hosts it and keeps it running. That is the whole role. We also build memory infrastructure, which means we compete in the thing we administer, so the limits are written down rather than promised:
- Our submissions go through the same approval as everyone else's. We do not merge our own.
- We do not decide who is admitted. The rules do.
- Our numbers are verified the same way as everyone else's.
- Wontopos publishes nothing on a new version for fourteen days.
- If other memory companies want to co-administer, that's better than us alone, and the offer is open.
A few of the submission rules point the same way:
- Publish the per-question record. Anyone can recompute the number from it. The aggregate is a claim, the record is the evidence.
- The reader, the judge and the prompts are set by the version. A submitter does not choose them.
- The harness has to be one a customer could use. A number produced through a path only its author can reach is not a number anyone else can get.
Everything there is Apache 2.0, so if we ever become the problem, the whole thing can be taken and run elsewhere without asking us.
It's all here: github.com/wontopos/glasshouse. Fair warning about what you'll find. The benchmark itself isn't in there yet and submissions/ is empty. The rules went up first, and we haven't submitted either.
The corpus
One person's life, told across about 17 months of conversation, with the facts you're supposed to remember buried inside it. 103,572 turns, 1,991 sessions, roughly 1.9M tokens.
It runs at four haystack sizes, from a small core up to the whole thing. The 1,882 turns that actually contain the answers are identical in all four, character for character. What changes is how much unrelated conversation is packed around them, so you can watch a system degrade as the haystack grows instead of getting one score and no idea what it means.
The axes
1,547 questions across 14 axes at the full size. Most are ordinary recall: it was said, can you get it back, can you get it when the question uses different words, can you say when it happened, can you combine two facts.
These four are why I built this.
Stale facts. A value changed. The new one exists, but I made it hard to find on purpose. Three outcomes instead of two:
| Answer | Score |
|---|---|
| New value | 1.0 |
| "I don't know" | 0.5 |
| Old value, stated as current | 0.0 |
The reason for the middle row is the whole point. Put "I don't know" and "confidently out of date" in the same bucket and you've hidden the thing that actually hurts in production, because those are very different things to be paying for.
Contradictions. The conversation states two different values for the same fact, nobody corrects it, and nothing tells you which is right. There is no correct answer. Confidently picking one is wrong. Saying "these don't match" is right.
Apparent contradictions. The mirror image. Two statements look like they clash but hold under different conditions, so both are true. Telling this apart from a real contradiction is the point.
Abstaining. Questions about things never mentioned at all. An empty answer scores full marks, anything invented scores zero. On top of that there are 500 false-memory probes, which plant something that was never said inside the question itself and check whether the system plays along.
Images
50 photos are shared inside the conversation, with 150 questions about them. Every one of those questions is pinned to a date: "In the photo from 14 April 2026, what was on the sofa?"
That's there because with 50 photos, a question like "was the laptop open?" points at two of them at once. Inside the flow of the conversation that's fine, but a question lifted out on its own, or translated into another language, becomes impossible to answer correctly. The date narrows it to exactly one photo, and it doesn't leak anything, since nothing is asking when it happened.
Languages
The corpus carries 100 sessions in 10 languages, 10 sessions per language, mixed in with everything else. 800 questions run on this, in both directions:
- A fact stored only in another language, asked in English.
- A fact stored in English, asked in another language.
Both directions matter because they break differently. One tests whether anything crosses the language boundary at all, the other tests whether the query side does.
Speed, and why it's measured this way
Speed matters, but measuring it fairly is harder than it looks, and this is the part I rewrote the most.
The problem is geography. Raw latency includes the speed of light. Measure from Seoul against a server in the US and you're 190ms in before any work has happened. Report that as-is and you're ranking where the server sits, not how good it is.
So the network floor gets measured separately and subtracted. What's left I call "latency minus round trip", not "pure compute", because response transfer doesn't fully subtract and calling it pure compute would be overstating it.
Two conditions reduce the leftover error, and both are part of the procedure:
- Warm the connection first. TCP starts slow and ramps up. Without warming, a response crossing 14KB picks up an extra round trip, and a 3,700 token response sits right on that boundary.
- Record the response size. The remaining error scales with it, so writing it down lets a reader judge how much slack is in the number.
The budget. People stop feeling like a conversation is flowing at around 1,000ms. The LLM's first token eats about 500ms of that on its own. So memory gets the remaining 500ms, and that's the target. Twice as fast as the target scores +1, four times slower scores -1, and it's a log scale in between, because the difference between 200ms and 400ms matters more than the difference between 3s and 3.2s.
Per question, and capped. Each question scores between -1 and +1, and the results are averaged rather than summed. Sum penalties instead and the total sinks on its own as you add questions, which means the score stops describing the system at all.
Speed is reported next to accuracy, not folded into it silently, and the weight is stated. And if speed can't be measured on a given system, the term is removed rather than zeroed. Being unmeasurable shouldn't be a penalty.
Nothing has been scored yet, and I'm not quoting numbers before there are numbers.
One question for you
If you were going to run something like this, what would have to be in it before you believed the result? And if you've ever looked at a published memory benchmark score and thought "no", what tipped you off?
Top comments (19)
You already have the number that answers your own question, in the record you linked. I downloaded the eight LongMemEval runs from
beam1m-tablet-2, reproduced all eight scores exactly (95.2 / 96.0 / 96.0 and 93.2 / 94.0 / 93.8, with theok: nullconvention leaving one question out of the three GPT denominators), and then computed the thing the paper can't compute from the aggregate: the per-question pairing.Reader swap vs. a plain re-run, per question. Over the nine opus↔gpt run pairs, the two readers disagree on about 27 of 500 questions (4.8–6.4%). Opus wins 171 of those and loses 75, so the direction is solid: +2.14pp, se 0.35pp, z = 6.1. But now the same statistic for two runs of the same reader: 12–26 questions flip, about 19 of 500 (2.4–5.2%). So the reader effect is roughly 1.4× the run-to-run noise, and a single pair of runs is what a leaderboard row actually is.
That gives the number a version is missing — its resolution, published next to the reader and the judge:
The per-pair standard deviation of a gap is 1.05pp, so anything closer than ~2pp at one run per system is inside the noise. Your own two LongMemEval rows for tablet-2 differ by exactly 2.0pp (95.7 vs 93.7) — the row you use to make the reader visible sits on the resolution line. Compute this once per version (re-grade a fixed slice with a second reader, pair it per question) and "five is certified" starts to mean something specific: certified above the resolution, not merely five runs.
And the band is not the reader's alone, which changes how it can be quoted. Conditioning on your own
gold_in_poolfield, on the ~376 questions where the engine delivered the evidence, the reader swap is worth +0.8 to +1.9pp (3.5–4.8% of questions flip). On the ~123 where it did not, it is +3.3 to +6.5pp (6.5–11.3% flip). The reader effect lives almost entirely in the questions where retrieval failed and the reader fell back on its own priors. So it is an interaction — reader × your engine's failure rate — and it will shrink for a better engine and grow for a worse one. "That gap belongs to the reader" is the right arc in the claims table, but a swap number measured on tablet-2 will not transport to a competitor's submission. Two ways out, both cheap: publish the band per submission, or publish the gold_in_pool split, which you already record.Three things I'd want before believing a result, since you asked:
glasshouse-v0.1/README.md— the reader and judge are named, but how much either of them moves a fixed set of answers is not.headline.scoreandheadline.stdev— aggregate against aggregate, which throws the pairing away and is roughly twice as wide as it needs to be. Comparing two systems on the identical question set, counting the discordant questions and stating the gap in questions rather than points is a two-line addition to whatever reports the headline, and it is the difference between "1.2pp apart" and "8 questions apart".The limitation on my numbers, stated rather than left implied: the two reader groups are different runs, so my pairing controls item difficulty but cannot separate the reader from the run draw — that separation would need the same answers graded twice. That is not a criticism of the record, it is why I compared the swap against the same-reader re-runs, which is the piece of your design that made it measurable at all.
To your second question: the tip-off is a table where one system has three numbers and none of them carries a resolution — the LoCoMo row in your own README (84 / 58.44 / 75.14) is that table. Not dishonesty, just a benchmark whose noise floor was never measured, so every reading had to be argued instead of settled.
Want this as a PR under
individuals/? Your own README says a recomputation is evidence rather than a proposal, and it should go to the submission it disputes — so tell me which of the two you'd prefer and I'll put it where it belongs. The script that produced the table is a couple of dozen lines and I'm happy to hand it over either way.Thanks, this is really useful. Glasshouse v1 isn't in the repo yet, so there's nowhere for a PR to land right now. It's going up soon. I'll message you when it's there, and if you send the PR then I'll review it and update.
Good — and the wait is fine, I would rather send something checkable than something fast. When v1 lands I'll bring the run I already did against your published
beam1m-tablet-2records (the eight LongMemEval files): reproduce all eight headline scores, then the per-question pairing — reader-swap flips against same-reader re-run flips, and the smallest gap each run count can settle. That was the substance of my comment, so it should go where the number it disputes lives.One piece of it does not need v1, because it is a property of the version rather than of any run.
glasshouse-v0.1/README.mdnames the reader and the judge; what is missing beside them is how much either one moves a fixed set of answers. That is computable once per version — re-grade one fixed slice with a second reader and pair it per question — and it is the kind of number that is cheapest to write down before any submission exists, since every later row inherits it. Rule 8 says five runs is certified; a resolution figure is what would make "certified" mean something specific, and it would let a reader tell a 2pp gap at one run from a 2pp gap at five.If a spec sentence is easier than a per-version computation, the equivalent is to require submitters to report their discordant counts against a fixed reference run: you get the same number out of the submissions, and it costs you nothing to host.
Not expecting either in v1 — the version is drafted by companies and rule 10 sends additions to the next vote — just flagging what I would put in the PR so you know what is coming.
It's up. github.com/wontopos/glasshouse, under glasshouse-v0.1. Said I'd let you know.
The complaints list you started from is the real spec: published vendor numbers don't match measured ones, and swapping the grading model moves results more than the gap between the systems being compared - which makes most published numbers decorative. Writing the limits in the repo instead of promising them is the only credible move when you compete in the thing you administer. "We do not merge our own" is the sentence that matters, because it converts trust-me into check-me. The grader-model sensitivity is the sleeper issue in the whole space: if the judge choice dominates the system choice, the benchmark measures the judge.
You compressed my whole fairness section into one line, and the line is better
than the section. On grader sensitivity I should be straight: I don't measure it. There's a check that forces the string scorer and the LLM judge to agree (it caught an axis
where "I don't know" scored 0.5 one way and 0.0 the other), but that's not the same as swapping the judge and seeing how much moves. All I have is the per-question record, so someone can re-score with their own judge and publish the gap. Preferably someone who isn't me.
v0.1 is out: github.com/wontopos/glasshouse
The three-score stale fact axis is the bit I'd steal. Every binary eval I've touched mixes "no idea" with "confidently out of date" and you just can't see what's actually breaking in prod. The geography correction on latency: we ran a comparison this year where the Seoul-to-US round trip was 190ms and the gap between the systems was 40ms, so the whole result was just geography. Worth adding at some point: distributed evidence, where the right answer isn't in one turn but spread across five, because that's quietly where a lot of memory systems fall apart.
40ms against a 190ms trip is a better version of my own point. Thanks for that. You're right about distributed evidence. I counted: 3 questions out of 1,547 have the answer spread across five turns, and the gather axis is all two-turn. So it's basically not in there. I'll look at adding it myself. The benchmark isn't in the repo yet though, so once I've put the draft up, a PR from you on this would be appreciated.
v0.1 is out: github.com/wontopos/glasshouse
Publishing per-question records helps reproducibility, but the conflict of interest moves upstream: who decides which memory failures enter the 14 axes and how much each axis counts?
A benchmark can be reproducible and still favor the maintainer’s product strengths. I would want versioned rationale for every axis and weight, blind proposals from competing vendors, and a holdout set controlled outside the host.
Which part of the benchmark definition can a co-administrator actually veto?
Fair, and the weighting part I can answer now. Every axis has a written rationale behind its weight, including the one criterion I used to sort them and the two weighting schemes I tried and dropped. It isn't in the repo yet because what's up there is still a partial draft. When the full draft goes up that document goes with it, so as of right now you're correct that there is nothing for you to check. Blind proposals and an externally held holdout are both good and I hadn't considered either one. I'll look at both properly. On the veto question, you're right that there's no answer in the repo. That needs defining rather than hand waving, and I'd rather write it down carefully than give you something off the cuff here. Thanks for all three. If any of it ends up in the benchmark I'll come back to this thread and say so.
That sounds like the right sequence. Publishing the rationales and the weighting schemes you rejected makes the trade-offs inspectable; blind proposals and an externally held holdout reduce the ability to tune the benchmark toward the product. The remaining piece I’d put in the same governance note is change authority: who may alter axes or weights, what evidence is required, and who can veto a change after seeing results. Looking forward to the full draft.
The fourteen-day rule is the one I would look at again, because as written it hands you the advantage it is there to remove. Everyone gets a new version at the same moment, but a competitor has every reason to publish early, and the rule guarantees you fourteen days of iterating against a version whose leaderboard is already carrying other people's locked-in numbers. A submitter who publishes on day two cannot see what you are aiming at. You can see exactly what you need to clear.
The fix lives in the same family as your other rules rather than needing new machinery: hold every number on a new version for the same fourteen days and release them together, so nobody gets to tune against a published bar. The asymmetry right now is not that you publish late, it is that you are the only party who can choose to.
Good catch, and you're right that as written it works against itself. A submitter who publishes on day two locks their number in, and I get two weeks of looking at it before I have to commit to mine. That's the opposite of what the rule was for. I'll look at the rule again. If it changes, I'll come back to this thread and say what it became. Thanks for reading it closely enough to find that.
v0.1 is out, in case you want a look: github.com/wontopos/glasshouse
Building a benchmark for your own product is always a conflict of interest. The best approach I have seen is using a third-party eval suite as an independent sanity check.
v0.1 is out, and the weighting document went up with it: github.com/wontopos/glasshouse
Some comments may only be visible to logged-in visitors. Sign in to view all comments.