We built a 30-question exam that tests whether an AI organization can correctly remember 130 days of its own operating history.
Before publishing it, we sat our own stack down and took it ourselves. First attempt: 11 out of 18. Then we audited all seven failures — and every single one was the grader's fault, not the examinee's. Two questions had wrong answer keys.
That's the thesis of this post: the most dangerous failures in AI agent stacks live in the verification layer, not the models.
The Three-Arm Experiment
Same backbone (MiniMax-M3), same questions, only the memory path changed:
| Arm | Score |
|---|---|
| Direct access | 16/18 |
| Through mem0 (default, semantic compression) | 11/18 |
| Through mem0 (compression off) | 18/18 |
One config flag = 7 points. Write-time compression drops machine-checkable fields (artifact: null) that audits need.
TypeSafe Recompute
We also recompute TypeSafe's "444.6x cheaper" claim with 90 timed calls:
- vs opus 5: 440x ✓ (real)
- vs fast models (MiniMax, glm): 16-18x
The multiplier landscape is two clusters, not a spectrum. Fast cheap models all land in a narrow 16-18x band; heavy generators sit at 184-440x. A single headline number is marketing; the multiplier-vs-comparison curve is the information.
Four Protocol Primitives
- Three-state grading (PASS/FAIL/U, U never counts)
- Public grader (anyone can audit our implementation)
- Signed receipts (ed25519, results can't be silently edited)
- Challengeable verdicts (90-day window)
Full artifacts: github.com/chunxiaoxx/nautilus-compass
Top comments (0)