DEV Community

Cover image for Every AI coding agent tracker is a self-report system

Every AI coding agent tracker is a self-report system

Alberto Clemente on August 13, 2026

On 27 July I opened a project I'd been building with Claude Code and found three things true at once: a card had carried a null commit for two da...
Collapse
 
reidmarlow profile image
Reid Marlow

The part I like is making the check an argv array owned by config. That sounds small, but it removes the agent's easiest escape hatch. I would still want the tracker to store the raw command output too, because a green exit code without the transcript is hard to debug later.

Collapse
 
albertoclemente profile image
Alberto Clemente

Failures keep the output — exit code, duration, head and tail of the log with the
elision stated. Passes don't: just the check name, argv, exit, duration, sha, and
whether the tree was dirty.

The reason is that the note is the memory. Every future session re-reads it, so a
full test log is a cost paid forever, not just by the run that produced it. And a
green run at a clean sha you can re-derive — check it out, run the same argv, the
transcript comes back.

The dirty tree is where you're right. A pass over uncommitted changes isn't
reproducible from the sha, so there the transcript was the only copy. That I
should fix.

Collapse
 
albertoclemente profile image
Alberto Clemente

Dirty passes keep the output now — 9feb4bd.

Clean ones still don't, and I'm leaving that: a clean sha can be checked out and re-run to get the log back, so storing it forever costs every future session for nothing. A dirty tree can't be re-run, which is exactly the case you spotted.

Collapse
 
peterbuildssecure profile image
Peter

The machine-owned check is the right boundary. I’d also bind every successful check result to the exact Git tree it evaluated.

An exit code alone can become stale immediately: the agent can run the check, modify another file, then hand the card back using a result produced against different state. Recording the commit/tree hash, command, exit code and output digest makes the claim reproducible.

For dirty worktrees, either refuse completion or hash the relevant files before and after the check and fail if they changed. The useful invariant is not just “this command passed,” but “this command passed against the exact artifact now being marked complete.”

Collapse
 
albertoclemente profile image
Alberto Clemente

"This command passed against the exact artifact now being marked complete" — that's the invariant, said better than I've managed it.

Most of it's there: argv, exit code, duration, sha, dirty flag, anchored to the tree
the check ran on rather than HEAD at write time. And the agent never carries a result at all — done() runs the check itself.

But you've found something I hadn't. I read the head before the run and never again, so a tree that changes during the check is recorded against pre-run state. Your before-and-after hash is the fix, and I don't have it.

Collapse
 
albertoclemente profile image
Alberto Clemente

Fixed in 9feb4bd. You were right — the tree was only read before the check ran, so a pass could be recorded against a tree that had already changed. It's read on both sides now, and if it moved the card doesn't move either.

One thing your comment saved me from: it hashes the diff, not the filenames. A file edited twice with the same name would have slipped through otherwise.

Thanks. First bug on that board that didn't come from me or CI.

Collapse
 
peterbuildssecure profile image
Peter

Nice fix. Hashing the diff contents closes the “same filenames, different state” hole.

The remaining edge is the small TOCTOU window between the second tree read and the card-state write. If another process can modify the worktree there, I’d make the final transition a compare-and-swap: update the card only if the current tree still equals the verified hash. Alternatively, hold the same repository lock across verification and transition.

Then the invariant becomes atomic: either the exact verified tree is marked complete, or nothing moves.

Thread Thread
 
albertoclemente profile image
Alberto Clemente

Fixed in 13450bb — the promotion is a compare-and-swap now. Of your two fixes it was the only one this design could take: the check deliberately runs outside the write lock (waiters give up at 60s, and a suite can run for minutes), so holding the lock across verification was out. Instead done() takes a third tree reading while it holds the lock, and the card moves only if the digest still equals the one the check was verified against. A mismatch lands in the same bucket as the last fix — absence of evidence, neither pass nor fail, the card stays in progress. The test stages your scenario literally: a second process holds the lock, the tree moves while the write waits on it, and the card must not move with it. Which is your closing sentence as an invariant: either the exact verified tree is marked complete, or nothing moves. Three findings now, each a strictly smaller window than the last. The acknowledgements section is developing a habit of your name.

Thread Thread
 
albertoclemente profile image
Alberto Clemente

@peterbuildssecure Postscript, a week on: all three of your findings shipped, and it's on npm now — npx shipward setup ~/code/your-project --seed-from-branches — so the thing a stranger installs today is meaningfully more correct than the thing I wrote about. That's down to you. Three rounds, each window smaller than the last, and I hadn't seen any of them coming.

If you ever have an idle hour, the compare-and-swap is the part I'd most like broken by someone who isn't me. You've got a better record at that than I do.

Collapse
 
mnemehq profile image
Theo Valmis

Who has the authority to assert this is exactly the right question, and it generalizes past task trackers. The same gap exists for architectural claims: an agent says this follows the pattern we agreed on, and unless something with actual authority checks that against a rule, you've got a self-report system for architecture too, not just for task status.

Collapse
 
albertoclemente profile image
Alberto Clemente

Agreed, and task status is actually the easy case — git doesn't care what the
agent claims, the commits and passing tests are either there or they're not.
Architecture is harder because most of the time the rule was never written
down in a checkable form in the first place. It's a paragraph in a design doc
or something agreed in a meeting, and you can't verify a claim against that.
Tools like dependency-cruiser or ArchUnit cover the parts you can encode
(imports, layer boundaries), but there's always a portion that won't encode.
I think the honest position is that this portion stays self-reported, and
the job is to keep it small.

Collapse
 
nazar-boyko profile image
Nazar Boyko

If the agent can select which declared check to run, what stops it from picking the cheapest one? A card handed back against the lint check instead of the test suite would still get the green status. Curious whether checks are bound to specific cards or any declared check satisfies any card.

Collapse
 
albertoclemente profile image
Alberto Clemente

Nothing stopped it, and the honest answer as of this morning was: recorded, not gated. Any declared check satisfied any card — a done() against lint earned the same green as the suite, and the only defense was that the evidence names its exam. You also found something sharper than your own question: a selected check became the card's check, so one cheap hand-back quietly lowered the standard for every later one.

Fixed in e10c1da. The rule is now that selection fills a blank, it does not change a standard: a card that already carries a check refuses a hand-back naming a different one — nothing runs, nothing is proved, same bucket as a failing check — and switching takes force:true, which writes a decision entry naming both checks, even when the chosen check then passes. There's no strength ordering between commands, so the tool can't know lint is weaker than the suite; what it can hold is that a standard never changes silently. The declared-by-human rule closed "the agent writes the exam" — you're the third reader to find a hole like this within a day of looking, and this one closed "the agent picks the easiest exam". Thanks.

Collapse
 
albertoclemente profile image
Alberto Clemente

Postscript: e10c1da is on npm now, so the hole you found is closed in the version anyone installs — npx shipward setup ~/code/your-project --seed-from-branches.

Worth saying what your comment actually did. You asked a question instead of reporting a bug, and the question turned out to contain a worse problem than the one you were asking — I'd never have gone looking at check selection on my own. If you do point it at a repo and try to pick the cheap check, I'd like to know whether the refusal reads as obvious or as the tool being obstructive. That's the part I can't judge from inside.

Collapse
 
nitishkumarpro profile image
Nitish Kumar

The key insight here is that agent state should be treated as derived data, not trusted self-reported data.

This is a much more important distinction than “better prompts for coding agents.” The agent can be perfectly capable of producing correct code while still producing an incorrect representation of project state.

I particularly like the distinction between storage and arbitration. Putting the board in Git gives you history, but it doesn't make the board truthful. Letting authoritative signals like Git ancestry and independently defined checks correct the board is what actually closes the loop.

There’s a similar pattern in production systems: never let the component responsible for performing an operation be the sole authority for declaring that operation successful. The interesting engineering question is always, “what independent evidence can prove this claim?”

That also makes the shell: false failure a useful example. The security boundary wasn't the execution primitive; it was who controlled the command definition. The same principle applies to agent tool permissions, CI pipelines, and automation systems.

Agents should be allowed to propose state. Systems should be responsible for proving it.

Collapse
 
albertoclemente profile image
Alberto Clemente

Six days is too long to leave this sitting — sorry, it deserved a faster answer than it got.

"Agents should be allowed to propose state. Systems should be responsible for proving it." I spent a month trying to get that into a title and you did it in two sentences. I'd have led the post with it.

Since you took the derived-data framing further than I did, here's where I found that it stops. Derivation covers what git can prove: a card claiming it shipped when the commit isn't an ancestor of main, a note that was true fourteen commits ago. The audit corrects those without being asked, and only ever forwards — it fills blanks, confirms what landed, and never moves a card backwards. What it can't derive is intent. No commit records that something is backlog rather than review, or that it matters more than the card beside it. That part stays proposed, and the system has nothing honest to say about it. Keeping those two categories apart is most of the design.

Your shell: false point also had a second half I hadn't seen, and a reader found it two days after the post went up. Owning the command definition stopped the agent writing its own exam — but any declared check still satisfied any card, so it could hand back against lint instead of the suite and collect the same green, and that choice then quietly became the card's standard for every hand-back after it. Fixed in e10c1da: selection fills a blank, it does not change a standard. Your principle, one layer down — the definition was owned, the choice wasn't. You'd probably have got there faster than I did.

Collapse
 
zira125 profile image
Zira

The authority-owner framing maps directly to agent deployment too. A worker should not be allowed to both perform a tool call and declare its outcome: the durable ledger or provider response has to establish that claim. I use a drain boundary where new claims stop, in-flight effects become terminal or UNKNOWN, and a replacement can reconcile from the idempotency key. Otherwise a self-reported “done” can hide the same gap as a stale tracker card.

Collapse
 
albertoclemente profile image
Alberto Clemente

The state you call UNKNOWN is the one this tool had to learn to say out loud. A check that times out, or runs while the tree moves under it, is recorded as exactly that — absence of evidence, neither pass nor fail — because collapsing UNKNOWN into either direction is where the lying starts. Earlier today my own machine made the case for me: a neighboring build starved the test suite past its time budget, and the gate refused its own project's hand-back twice rather than call a timeout a verdict. Your drain boundary is the same invariant one layer up: nothing may claim an outcome the durable record can't establish. The ledger outranks the worker, git outranks the board.

Collapse
 
glenallen profile image
Glen Allen

This makes me think of agent tracking less as a logging problem and more as an authority-design problem. If the same system can perform an action, declare it successful, and define the criteria for success, the feedback loop is inherently biased. Independent sources of truth become much more valuable as agent autonomy increases.

Collapse
 
albertoclemente profile image
Alberto Clemente

That's the design position, yes — and the one refinement I'd make is that the fix isn't only independent sources of truth, it's removing the agent from the reporting path entirely. The agent here doesn't run the check and carry back the result; it can only request the transition, and the tracker runs the check itself, under the same lock that moves the card. An agent that never holds the evidence can't bias it. The criteria are out of its hands too — a check is declared by the human, as an argv array in config, and a card can only name one, never define one. What's left to the agent is the part you'd actually want an agent for: doing the work and saying what it thinks it did.

Collapse
 
glenallen profile image
Glen Allen

The “who has the authority to assert this?” framing is a really useful way to think about agent reliability. In our AI work at IT Path Solutions, we’ve found that different claims often need different sources of evidence and different validation windows. A Git commit can confirm that code landed, for example, but it doesn't establish that the behavior remains correct after dependencies or surrounding logic change. Giving each claim a clear evidence owner, scope, and freshness boundary could make these systems much more trustworthy.

Collapse
 
albertoclemente profile image
Alberto Clemente

The freshness boundary is the part of this the tracker actually has an answer for. Every evidence entry is stamped with the sha it ran against, and every time a session re-reads the board the drift since that sha is computed and surfaced — commits since, and whether any of them touched the files the entry names. So a green check doesn't quietly stay green: it visibly ages. What it can't do is tell you which later changes invalidate it — that would need the dependency knowledge you're describing, and I don't think a tracker should pretend to have it. "This passed at that sha, and 14 commits have landed since, 2 touching files it names" is a claim git can back. "It still holds" isn't, and the honest move is to keep those two sentences apart.