DEV Community

openamer
openamer

Posted on

Most of our agent's runs fail. The ledger records every one of them.

Most of our agent's runs fail. The ledger records every one of them.

OpenAmer is a self-verifying AI agent that runs on a CPU-only Windows laptop. It keeps a public
outcome ledger — memory/si/outcome_ledger.jsonl — where every run it takes is recorded as a
row with a timestamp, the task, the goal, the agent that ran, and a passed: true/false verdict.

The part nobody expects: the ledger does not hide the failures. It is bigger than the successes.

The numbers, measured today

Metric Value
Total recorded outcomes 1,701
Passed 543
Failed 1,158
Pass rate 31.9%
First entry 2026-09-28
Last entry 2026-10-11

That pass rate is not a bug to be embarrassed about — it is the whole point of the design. A system
that only records successes cannot tell you why something went wrong, and "why" is the only thing
that makes the next run better.

What each failure row actually contains

Every row carries a why field. A failure is not just passed: false; it is a failure with a
reason:

  • the task timed out
  • the tool returned an unexpected shape
  • the agent took a path that did not reach the goal
  • a dependency was missing at the moment of execution

The ledger is written before the side effect, and the side effect is only kept if the ledger row
is durable. That ordering is the anti-double-charge mechanism: an effect can land while its receipt
is missing, but a receipt can never exist for an effect that did not land. "Effect landed, receipt
missing" is itself a recorded terminal state, not a silent hole.

Why we publish the failures

Three reasons, in order of importance:

  1. Self-measurement without a failure record is just optimism. A dashboard that shows only green is a decoration. The ledger's value is that the red rows are real and countable.
  2. The failure distribution is the training signal. The self-improvement loop reads the ledger, and the failures dominate it. Improving the pass rate means working on the 68% that fail, not congratulating the 32% that pass.
  3. Anyone can verify it. The ledger file is in the repo's memory directory. You can count the rows yourself. There is no marketing claim here that cannot be checked in one command.

What this is not

This is not a claim that OpenAmer is the best agent in the world. It is not even a claim that it
is good. It is a claim of a specific, narrow, verifiable kind: the agent can tell you exactly how
often it fails, and why, in writing, in a file you can read right now.

That is a lower bar than "best in the world," and it is the bar we actually meet.

The repo

github.com/openamer/openamer — Apache-2.0, CPU-only, Windows-native, self-improving.

The ledger is at memory/si/outcome_ledger.jsonl. The numbers above were read from it directly;
they change as the agent keeps running, so check them again before quoting them.

Top comments (0)