DEV Community

Cover image for The Case for Evidence Based Agentic Development
Kent Alstad
Kent Alstad

Posted on AI-assisted

The Case for Evidence Based Agentic Development

This is evidence that does not forget. It is not confused by the excitement of the moment. It is not absent because human witnesses are. It is factual evidence. Physical evidence cannot be wrong, it cannot perjure itself, it cannot be wholly absent. Only human failure to find it, study and understand it, can diminish its value.

Paul L. Kirk, 1953 [1]

Kirk was writing about crime scenes. I think he was also writing about software built by agents.

Where we are

Software is more complicated and more feature rich than it has ever been. Agents did that. Point enough of them at a well-defined problem and you get working code, passing tests, and a system that looks almost done, in days.

What we don't have is a viable way to stabilize it and release it.

I found this out the way most people will. I realized I could run five projects at once. Five repos, five sets of agents, five streams of change, all day. It worked, in the sense that code appeared. It also became impossible to monitor every session. I could not read what was being produced, let alone judge it.

That was the moment the framing changed for me. I don't want a co-pilot. I want five dev teams. And a dev team needs something a co-pilot never did: a way for the person in charge to know what is true about the product without having watched every line get written.

The phases of a product

Every product I have built with agents goes through the same three phases, and the proportions are consistent enough that I'll put numbers on them.

The first 80% comes from the spec. You describe what you want, the agents build it. This part is astonishing and it is the part everybody has seen.

The next 15% is vibing. Feedback, iteration, trying things. You use the app, you say "no, like this", the agent changes it. This is where most developers are when they are rapidly creating something. It is interactive, it is the closest to the old way we developed, and it is great for creative solutions. It is also where most people stop, because the app looks done.

The last 5% is the hard part. Bugfixes. Performance. Hardening. Integration across the pieces, and stabilization before release. The trick in this phase is not moving backwards, and not injecting too much change while you fix what is in front of you. Of the five phases a release goes through, integration and stabilize are the difficult ones, and they are the two the first 95% does nothing to prepare you for.

The phases of how people approach it

People's approach to agentic development goes through phases too. I went through all of them.

First, the pair model. For a long time, the pairing model that tools like Cursor promote seemed right. One developer, one agent, a conversation, a diff. But that model is steeped in helping us do it the way we used to do it, only faster. That doesn't work. You end up with more code to stabilize and no reasonable strategy to do it. The pair model is the 15% phase made permanent. It has a place. It is not a viable solution for stabilization, and it does not scale to factory size.

Then, more tests. The next phase was to make the agents write tests for everything, and the problem became how to manage a large volume of change in a way where the tests really guard against going backwards. For the most part, the agent that writes the code also writes the tests. They helped, for APIs. But even there, I found that the complex ways the apps actually used those APIs let problems through that the API tests passed. And what the tests did not catch at all was the requirements mistakes. A test written by the author of the code checks the author's understanding of the requirement against itself. If the understanding is wrong, the test is wrong in the same way, and it passes.

Then, human review. It really seemed like human review was essential. It also seemed impossible, because the volume of review is overwhelming. Five teams produce more pull requests in a day than a person can read in a week.

So that is the situation. The volume and pace of change demands an oversight system that can keep up with "super" developer productivity. Not a faster reviewer. A different kind of gate.

How do we know it is done?

Before anything else, there has to be an answer to that question that was written down before the work started.

For every change, an acceptance plan: the criteria that define done, and the evidence items that will prove each one. If the change has a user interface, the evidence has to narrate the workflow that must be demonstrated. Not "the login works". The fresh install, the one button on the launch screen, the tap, the page the app must open, the consent screen, the picker that appears after it.

I'll describe what these plans look like in practice, because the details are the method.

  • The criteria are numbered, and each one is a testable statement about behaviour, not about code.
  • Each evidence item names the criteria it covers and the artifact it produces: a log, a test report, a recording, a screenshot.
  • Each item says who produces it, and the answer is never the agent that did the work. The wording in ours is blunt: produced only by the evidence service, after the build. No agent runs, copies or edits it.
  • For a user interface, the item walks the workflow step by step. Every step has a test id. Every step that has to be stubbed is printed into the run log as a stub, so the recording can't pretend.
  • Test data is a dedicated, human-provisioned test account. Never a production user.
  • Honest limits are written into the plan. If a step is interactive and cannot run unattended, the plan says so and says how the evidence will be produced instead.
  • A human accepts the plan before any code is written.

That last point matters more than it looks. The plan is accepted while nobody is attached to an implementation yet. It is the one moment in the process where "what does done mean" gets decided without a diff in the room.

The case: concrete evidence, human focused, end-user driven

Here is the argument.

Not only do we need concrete evidence to check each change against. The evidence has to be human focused and end-user driven. It has to show the product doing what a person would do with it.

It may seem counter-intuitive, but when an agent is tasked with writing a test fixture that works the feature the way a user would, and that fixture captures a video, and another agent uses that video as the working evidence, the agents start to catch errors they could not find in the code. No person records anything. The video is just test output. What makes it different from a test report is its form: it is meaningful to the human who has to accept it, and it turns out to be meaningful to the agent that compares it. The question "has this video changed compared to the baseline?" is a different question from "is this code correct?", and agents are better at the first one. The interpretive skills they bring to comparing two recordings yield a better level of oversight: more attuned to human requirements, and to whether the app is actually usable.

Two things I want to be clear about.

This is in addition to tests, not instead of them. I have thousands of unit tests and I am keeping every one. Videos are strong exactly where unit tests are weak: the whole product, in use, end to end. Unit tests are strong where videos are weak: a single function, every edge case, in milliseconds. Evidence is both.

The video proves more than what it shows. For the fixture to record a user signing in and picking a repository, the build has to succeed, the device has to boot, the network has to be up, the backend has to answer, the data has to be right, and the whole thing has to finish inside the timeout. If any subsystem the user never sees is broken, there is no video. Often the issues the recording exposes are not in the frames at all. They are in the fact that the recording could not be made.

How a recording is made, and what a baseline is

The acceptance plan lists the criteria and the evidence items. An agent writes a test for each item. A separate service, not the agent that wrote the code, builds the branch and runs the test. The test produces the output: a recording, a log, screenshots. A reviewer accepts or rejects.

Accepted output becomes the reference. That is the baseline, and this is the part I want people to take away. The baseline is produced when the evidence is accepted. Nobody records a baseline by hand. Nobody decides in advance what the product should look like. The last thing a human accepted is, by definition, what the product should look like now.

The next change to that area runs the same test, and the new output is compared with the reference. Because every change is issue driven, with an acceptance plan, and ends with evidence accepted, the agent knows what was supposed to change. So it can tell the difference between "this moved because the plan said it would" and "this moved and nothing said it should". It flags the evidence that needs to change as part of the process, and it flags the differences that nothing explains.

People accept the reference. After that, the agents compare against it.

Where the gates sit

Quality sets the schedule, not the calendar. A version is in one phase at a time, and it moves to the next when the evidence says it can.

The flow from acceptance plan to accepted evidence: plan and spec, agent writes the fixture and builds the change in isolation, the evidence service runs the fixture and produces the output, an agent compares it with the reference, a human reviews what passed, acceptance gates the merge, accepted output becomes the reference

  • Before any code. The acceptance plan, and a design spec reviewed by an Architect for the design and a Product Manager for the product. Nothing is built until both are approved.
  • During the build. Each change is worked in isolation, on its own branch, in its own worktree. The way out is evidence: an artifact produced from a build of the code under test. A pull request is not evidence.
  • At integration. One change at a time. The evidence is reviewed and accepted by someone who is not the implementing agent, and then the change is merged. Acceptance gates the merge. Where a change can only be proven from an assembled build, because it spans repositories, the order flips: merge first, produce the evidence from the integrated build, and acceptance gates the version leaving integration instead. Either way, nothing leaves integration without accepted evidence, and integration testing runs after every merge.
  • Acceptance creates the baseline. Evidence produced for review lives in a scratch area. Accepted evidence is committed as the reference. Recordings, logs and screenshots all take the same path.
  • Stabilize. Slow the rate of change. Only the bugs the release needs. A risky feature gets flagged off rather than fixed under pressure.
  • Released. Each change is confirmed in the released build before the version is closed. Hotfixes only after that.

Waivers exist. Sometimes a change does not need a recording, or does not need a plan. But a waiver is recorded, with a reason, by a person. It is never a silent skip.

What the humans do

The human still creates the specs and the product definition. That does not go away, and nothing in this method touches it.

In the stabilization process, the human's job has three parts.

  1. Accept the evidence plan up front. Decide what done means before the work starts.
  2. Review the evidence that is ready for acceptance, after it has passed the agent gates. The agents pre-flight the review. They run the comparison against the reference, sweep out most of the problems, and send them back to be fixed before a person ever sees them. What reaches the human is evidence that has already passed, with the agent's account of what it checked.
  3. Review the repo news: the agent's report of the differences it flagged for investigation, with the narrative detail that lets you dig into only what is needed. You open the three recordings that changed when nothing said they should. You do not open the forty that didn't.

There are two human roles in this. The one I just described is a Product Manager. The other is an Architect, for the technical issues: the design, the data, the things a recording cannot show. In general, every review is either a business review or a technical review, and the method routes each piece of evidence to the right one.

I am not stuck on where humans participate. I will continue inserting an agent into the human role, prioritizing or helping with review.

Why would this be true?

I came to this by running out of alternatives, not by reading papers. But the published work points the same way.

Judges, human and machine, are more reliable when comparing two things than when rating one thing on its own. Thurstone showed it for people in 1927, and the finding has held up [2]. Surveys of language models used as evaluators find the same: pairwise comparison is more consistent and agrees with human judgement more often than absolute scoring [3].

Asked to judge correctness from the code, models miss a lot. In one study a frontier model passed half of the incorrect implementations it was shown [4]. In a 2026 code-review benchmark, four of five models found zero performance bugs [5]. Telling the reviewer the change was bug-free cut its detection of real vulnerabilities by as much as 93 percent [6]. That last one is the agent saying "done".

Tests written by agents are weak oracles. In a study of agent-authored test patches, four in five had weak or no explicit assertion: the test ran the code and checked nothing [7]. And models favour their own output. The more a model can recognise its own work, the more it prefers it, even when told the work is its own and it isn't [8]. The author's tests and the author's review are the same judge.

None of that proves that comparing recordings of an app catches regressions. Nobody has published that study. It says the method is pointed in the direction the evidence already points. The proof is what the rest of this series is for.

What I am after

I am an agent product developer trying to find a way to ship great products. The other side of this is the struggle to stabilize, and not shipping. I have been on that side.

Value, for me, is a shipping product that people use. If this method yields that, without slowing down, and without introducing new costs or processes that can't be managed, then it is valuable. If it doesn't, it isn't, and I'll say so.

So I'll end where I started, with Kirk. Evidence that cannot perjure itself, that is not absent because witnesses are, that is only diminished by our failure to look at it. When your agent tells you it is done, what do you have in your hand?

Are others doing this? Is there a better way?


Notes and sources

  1. Paul L. Kirk, Crime Investigation: Physical Evidence and the Police Laboratory (New York: Interscience, 1953), ch. 1, p. 4. Often misattributed to Edmond Locard.2. L. L. Thurstone, "A Law of Comparative Judgment", Psychological Review 34 (1927); D. Laming, Human Judgment: The Eye of the Beholder (2004).
  2. "A Survey on LLM-as-a-Judge", arXiv:2411.15594; Li et al., "A survey on LLM-as-a-judge", ScienceDirect, 2025.
  3. "On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization", arXiv:2507.16587.
  4. "Bigger Isn't Always Better: A Comparative Evaluation of LLMs for Automated Code Review", arXiv:2606.15689.
  5. "Measuring and Exploiting Confirmation Bias in LLM-Assisted Security Code Review", arXiv:2603.18740.
  6. "All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code", arXiv:2606.18168.
  7. Panickssery, Bowman and Feng, "LLM Evaluators Recognize and Favor Their Own Generations", NeurIPS 2024, arXiv:2404.13076.

Research notes: do agents judge differently when comparing than when verifying?

Supports the claim:

  • Comparing beats rating, for models. Surveys of LLM-as-a-judge find pairwise comparison more consistent (less position bias) and better aligned with human judgement than absolute scoring (arXiv 2411.15594; ScienceDirect S2666675825004564).
  • Comparing beats rating, for humans. Thurstone (1927), Laming (2004): people are more reliable at relative than absolute judgement. Comparative judgement in assessment rests on this.
  • Models judging correctness from first principles do badly. GPT-4 passed 50% of wrong Java implementations as correct (arXiv 2507.16587). In a 2026 code-review benchmark four of five models had 0% recall on performance bugs (arXiv 2606.15689). Framing a change as "bug-free" cut vulnerability detection by 16 to 93% (arXiv 2603.18740).
  • Agent-written tests are weak oracles. 80.2% of agent-authored test patches had weak or no explicit oracle (arXiv 2606.18168, "All Smoke, No Alarm"). Agent-written testing is "a model-dependent process style rather than a dependable driver of success" (arXiv 2602.07900).
  • Models favour their own output. Self-preference bias rises with self-recognition; models preferred summaries merely labelled as their own (Panickssery, Bowman, Feng, NeurIPS 2024). The author's tests and the author's review are the same judge.
  • Vision models find non-crash functional bugs from screenshot sequences that crash-based tools cannot: VisionDroid found 29 new bugs on Google Play, 19 confirmed and fixed (Liu et al. 2024). Third-party reruns found far fewer (FuncDroid, GraphDroid), so setup matters.

Cuts against, or limits, the claim:

  • Pairwise judging is more vulnerable to distractors the judge happens to favour; absolute scoring with a rubric was more robust to that manipulation (arXiv 2504.14716).
  • Pairwise verdicts can be locally consistent and globally contradictory (arXiv 2602.16610).
  • Video difference captioning is still hard. ViDiC-1K (1,000 video pairs, 3,720 checklist items) shows "a significant performance gap in comparative description and difference perception" across 17 multimodal models (arXiv 2512.03405). Same-or-different on image pairs: 17.6% for an untuned 4B model (arXiv 2501.04670). Frontier models are better; nobody has published numbers for app recordings.
  • No published study measures multimodal models on UI regression from recordings against a baseline. It is a gap to fill.

Build Stuff 2026

Control Your Fate at the Metal
Wednesday 2 December, 15:00. 45 minutes.

If you're coming, say hello. If you're not yet, 2026SPEAKER_20 takes 20% off a ticket.

#BuildStuff15 #buildstuffconf #AgeOfAgency @buildstuffconf https://buildstuff.events

Top comments (0)