DEV Community

woochan
woochan

Posted on

I Made a Memory Benchmark as Fair as I Could. 15 Stars, 0 Runs. What's Wrong With It?

I published a memory benchmark six days ago. I spent most of the build trying to make it fair rather than trying to make it look good, and now I have no idea whether that worked.

Here is the whole state of it:

stars        15
forks        0
issues       0
submissions  0
Enter fullscreen mode Exit fullscreen mode

People star it. Nobody runs it. Which means the one thing I actually wanted, someone telling me a question is broken or a rule is wrong, has not happened once.

What it is, quickly

Glasshouse is a long-term memory benchmark. 2,847 questions over a conversation that runs to 1.97 million tokens, in 10 languages, with 50 photographs. Every axis is reported on its own and there is no single headline number, because a system can be great at one and useless at another. It goes past plain recall too. If a fact changed and the system cannot find the new value, "I don't know" scores better than confidently giving the old one. If two stored facts disagree and nothing settles it, saying they do not agree is the right answer.

The reader and the judge are fixed by the version, so a submitter does not get to pick the model that reads the answers. Every submission carries its per-question record, so anyone can recompute the number instead of trusting it.

What I think might be in the way

I can guess at some of it, and none of these are accidents, which is exactly why I want to know whether I got the tradeoff wrong.

  • You have to write the ingest yourself. I deliberately ship no runner for storing the corpus, because the only client I could write is for the memory system my company sells, and that does not belong in a benchmark anyone is meant to be able to win. It does mean step one is work.
  • Scoring costs your own money. The judge runs through OpenRouter on your key.
  • The standard run is 1.97 million tokens. That is not something you try on a whim.
  • The official reader is claude-opus-5.5. I fixed it because changing the reader changes every number, but it is not the cheap option.

What I'm asking

If you opened it and closed the tab, I would genuinely like to know at which line. And if you are the kind of person who would run something like this, what would have to be different for you to actually do it?

Repo: https://github.com/wontopos/glasshouse

Top comments (3)

Collapse
 
sinarezaei profile image
Sina Rezaei •

I think the interesting problem here may be less about the benchmark itself and more about the distance between “this is worth testing” and “I can actually test it.” The fairness decisions make sense, but every additional step adds friction: writing the ingest layer, paying for the judge, and committing to a 1.97M-token run.

Maybe the missing piece isn't changing the benchmark, but creating a cheap “first run” path.
If someone can go from zero to a small, representative result in a few minutes, they might be much more willing to invest in the full benchmark afterward.

Good benchmarks need rigor, but they also need an onboarding path.

Collapse
 
woochan profile image
woochan •

Good point. There's already a core tier and a dry mode for exactly this, and it never occurred to me to mention either of them in the post. Thanks for this, I'll work it in from now on.

Some comments have been hidden by the post's author - find out more