Disclosure: I'm affiliated with Belcore, a memory and context layer for LLM apps. This is a measurement post. The numbers, their limits, and one result that did not hold up are all below.
Why "tokens saved" is slippery
Chat and agent setups usually keep context by resending history on every call, so input tokens per call grow with the session. A memory layer replaces that with a selected context. How large the saving looks depends on what you compare, so here are the same runs counted three ways.
Setup
- Dataset: LongMemEval_S, 500 questions, about 109k tokens of prior conversation each (roughly 500 turns).
- Answer model: gpt-5. Fixed seed, frozen harness.
- Tokenizer for context counts: o200k_base.
- Runs: August 2026.
Three numbers
| What is counted | Tokens per call |
|---|---|
| Full transcript, uncut | 109,079 |
| Assembled memory context | 6,671 |
| Billed input per call (27-question subset, includes prompt and question) | 9,908 |
- Context tokens: -93.9%.
- Billed input on that subset: about -91%. The subset was hand-picked, so treat this one as indicative only.
What this does not show
- Nothing here says anything about answer quality. Accuracy is a separate metric and I'm not quoting one in this post.
- Cost per successful task, including retries and fixes, is the number that matters for a real bill. I haven't measured it yet.
- LongMemEval_S is a public benchmark, not production traffic.
A result that did not hold up
We tried a multi-pass retrieval step that re-checks and re-retrieves before answering. On 27 questions hand-picked for missing evidence it went from 16 to 20 correct. On a random sample of 103 the attributable effect was +2, inside a ±5.2 noise bar, with one case where extra evidence turned a correct aggregation answer wrong. It also cost about 2.9x the input tokens and 2.6x the latency on the subset. We are not shipping it as an improvement.
Two more limits: in the wrong answers we audited, 38% involved a gold label that was wrong or defensible either way, and n = 499 cannot resolve small effects.
Try it
There is a free measurement demo on the site: belcore.xyz
Questions about the method are welcome, especially the ones that make the numbers look worse.


Top comments (1)