DEV Community

openamer
openamer

Posted on

What it costs to run a self-verifying AI agent on a CPU-only laptop

Most agent frameworks assume a GPU and a server. OpenAmer runs its core loop on an ordinary
Windows laptop - CPU only, no CUDA, no cloud round-trip for the core - and the part that surprised
us was not the model. It was how the agent proves to itself that it actually did what it claimed.

This is a short field report on the outcome ledger, with numbers read from the running install.

The problem: an agent that says "done"

Ask an agent to do five things and it will happily report success on all five. Some of them
didn't happen. A tool returned an error string instead of raising, a retry silently double-wrote a
file, a step was skipped because a dependency looked satisfied. None of that shows up in a chat
transcript, and none of it is caught by a unit test, because the failure lives in an integration
boundary the unit test never crosses.

The uncomfortable part is that the agent cannot honestly report on its own work using the same
machinery it used to do the work. The report is a claim. Claims need evidence.

The ledger

OpenAmer appends one JSON line per completed task. The keys are deliberately boring:

{"ts": "...", "task": "...", "goal": "...", "agent": "...", "passed": false, "why": "..."}
Enter fullscreen mode Exit fullscreen mode

There is no separate "success" stream and "error" stream. Failures are rows like any other row,
with a why field that says what broke. That was a design decision, not laziness: if failures go
to a log and successes go to a dashboard, the dashboard is marketing and the log is where you
actually look. One append-only file keeps them in the same accounting.

Live numbers

Measured from the running install on 2026-10-10:

metric value
ledger rows 1,312
passed 388
failed 924
window 2026-09-28 -> 2026-10-10

Read that ratio twice. Roughly seven in ten recorded tasks failed. We are not going to hide
that behind a "success rate" chart, because the number is the point: the ledger is a failure
detector that happens to also record successes. An agent that reports 95% success on its own work
is usually measuring the wrong thing.

Why CPU-only

Constraint breeds honesty. Without a GPU we could not paper over latency with throughput, so the
loop had to be cheap and the verification had to be synchronous - the row is written before the
next step starts, not batched at the end. A verification step that runs in a batch after the fact
is a report; a verification step that gates the next action is a control.

It also means the whole thing runs on hardware you already own. No cluster, no spend meter
ticking while you sleep.

Background computer-use

The desktop layer drives a real Windows session - clicking, typing, reading windows - without
taking over the cursor or the focused window, so you can work while it works. That is an
engineering claim we test rather than assert: the test suite launches actions against a live
desktop and checks the effect, not the intent of the code.

What this is not

  • Not "best in the world" at anything. It is a young project (a handful of stars) and the ledger says so out loud.
  • Not a benchmark entry. The numbers above are one install's own tasks, not a leaderboard score.
  • Not finished. 924 failed rows is a to-do list, not a trophy.

Where to look

Code, ledger format and tests: https://github.com/openamer/openamer (Apache-2.0).

If you are building agents, the transferable idea is small and it costs almost nothing to copy:
write one row per task, include the failures in the same file, and make the row a gate instead of
a receipt. Your dashboards will get uglier and your system will get more honest.

Top comments (0)