DEV Community

openamer
openamer

Posted on

How we made an AI agent measure itself: the outcome ledger

Most "my agent is autonomous" posts show a happy-path demo. A demo tells you the agent
can finish a task. It tells you almost nothing about the rate at which it finishes,
or about the failures — which is the number you actually need if you're going to trust
it. So when I set out to build a self-verifying agent that runs 24/7 on a CPU-only
laptop, the first thing I built was not a planner. It was a ledger.

Repo: https://github.com/openamer/openamer

The idea: every task writes one verifiable line

The agent doesn't get to claim success. Every task it runs appends a single,
append-only line to an outcome ledger with a self-assessment and a verdict:

{"ts":"2026-10-09T21:59:34Z","task":"...","goal":"...","agent":"...","passed":false,"why":"exit=1 Traceback ... AssertionError"}
Enter fullscreen mode Exit fullscreen mode

The passed field is the verdict; why is the reasoning. The file is append-only —
failures are written to the same place as passes, so you cannot hide them by omission.
Three consequences fall out of that one design decision:

  1. The metric is a rate, not a vibe. "Of N attempted outcomes, X passed" is a number finance can audit. "The agent did a lot" is not.
  2. Regressions become visible as a trend. A falling pass-rate is a signal; a feeling that the agent "seems worse lately" is not.
  3. Failures are first-class. Half the ledger being red is information, not embarrassment — it's the input the agent uses to decide what to fix next.

Real numbers (measured, not illustrative)

As of this writing the ledger holds 1,266 rows, 379 passed, 887 failed — first entry
2026-09-28, last 2026-10-09. Yes, the fail rate is high. It's a machine that is
self-hosting a lot of its own experiments and reporting every one of them, including the
bad ones. The point is not that the number is pretty; the point is that it is there,
it is honest, and it moves.

Why a self-verifying swarm needs this first

A swarm that can't measure itself will optimise for the demo and drift in the dark. The
ledger is the ground truth the whole loop hangs off: the self-improvement pass reads it
to find recurring failure motifs, the pass-rate trend is the health signal, and any claim
the agent makes about itself has to reconcile against a number in the file. It's the
cheapest possible foundation for "self-verifying" — an append-only log, a boolean, and a
reason string.

It also forces honesty in the other direction. Writing this post, the rule I held myself
to was: never state a count I haven't just measured. The numbers above came from running
a two-line aggregate over the live ledger immediately before writing them. If they're
stale by the time you read this, that's fine — re-run it.

Running it on a CPU laptop

The other half of the constraint: this runs on one laptop, no GPU. The split that makes
that viable is to keep cognition and plumbing separate. A small local model handles
decisions — what to do next, which tool applies. Everything else (the tools themselves,
scheduling, and the verification in this post) is ordinary deterministic code that needs
no accelerator. The bottleneck at that point is latency, not capability, and you buy that
back by keeping prompts short and letting the deterministic layer carry the weight.

What to take from this

If you're building an agent you want to trust, the smallest useful thing you can add
today is not another tool. It's a ledger: one append-only line per task, a boolean
verdict, and a reason. Everything else — trends, regression detection, self-improvement,
honest claims — is downstream of that.

Repo: https://github.com/openamer/openamer

(Numbers in this post are live, measured from memory/si/outcome_ledger.jsonl at
publication time. No benchmark rankings or "best in the world" claims — just the ledger.)

Top comments (0)