DEV Community

Ícaro Galvão do Nascimento
Ícaro Galvão do Nascimento

Posted on Originally published at icaro0310.github.io

I built tools to verify what my AI coding agent actually did

AI coding agents are great at one thing that isn't writing code: asserting. "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic.

I'm a QA engineer who runs Devin daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.

This is the short version of what exists, why, and what I learned.

The gap: agents assert, QA verifies

The trigger was simple. An agent session reported "fixed, all green" — and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the claim and the state are different artifacts, and only one of them is evidence.

The insight that made everything else possible: agent CLIs already write structured telemetry locally — session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.

Five tools, one pipeline

The tools compose into a sequence: understand → verify → measure → control → judge.

devin-internals-spec — understand. Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer — everything downstream reads against it.

devin-qa-pack — verify. Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against tool_call_state, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.

devin-evals — measure. Once verification works, you can ask the harder question: how good is the agent on this kind of task? Golden tasks, rubric scoring, regression tracking across sessions.

devin-bridge — control. A policy gate between intent and execution — ACP-based control with a --devin-only mode that enforces "this session does exactly what it was scoped to do."

poordjaevin — judge. The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part — the judge went from ECE 0.170 to 0.071 through iterative refinement on real session data. (ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server — poordjaevin serve — so agents can consult it mid-flight.

What "local-first" actually buys you

Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.

Concretely:

  • devin-history — cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)
  • devin-metrics — telemetry aggregation for quality signals over time
  • devin-memory + devin-search + devin-graph — memory store, retrieval, and a knowledge graph over sessions, projects and decisions
  • devin-doctor — environment diagnostics across Windows/Linux
  • devin-backup — snapshot + verify + restore of agent state
  • devin-janitor + devin-redact — cleanup and secret-redaction so state can move safely
  • devin-office — the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry — and it makes for a great demo GIF.

Three lessons from building this

1. The agent's own telemetry is an untapped QA datasource. Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" — and verification is what makes delegation safe at scale.

2. Calibration > confidence. A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 → 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.

3. Devin-only mode matters more than integrations. Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP — all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.

Try it / tear it apart

Everything is MIT-licensed, cross-platform, and documented in English + PT-BR:

  • Profile/catalog: github.com/Icaro0310
  • Website: icaro0310.github.io
  • Start here: devin-qa-pack (the verifier) or poordjaevin (the calibrated judge — pip install poordjaevin / uv tool install poordjaevin)

Honest question for the comments: if you run AI agents on real work — Copilot, Devin, Cursor, Claude Code — how do you verify their claims today? Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.

Top comments (0)