DEV Community

hefty
hefty

Posted on

Your Coding Agent Needs a Checkpoint Before Its Next Edit

An agent changes a settings screen, runs a test, and asks to keep going. The transcript is still scrolling. You can see that something happened, but the next edit may start before anyone has looked at the current one.

A stop button and a live log help, but neither tells you whether this slice is ready to become the starting point for the next one.

I'd put a checkpoint there. Show the exact candidate, the checks that ran, what's still unknown, and a decision the runner actually obeys. This is a proposed workflow, not a feature claim about any product discussed below.

Freeze the slice, not the whole conversation

Consider a hypothetical request to add a "pause notifications" control. The first slice adds the toggle and saves the preference. The agent then wants to change the notification list UI.

A useful checkpoint would look roughly like this (illustrative, not a real run log):

Candidate     <revision under review>
Scope         Pause toggle and preference save; list UI untouched
Check        <named test for saving the preference> — pass on this candidate
Browser      <named toggle-and-reload scenario> — not run
Open gap     Keyboard and screen-reader behavior not reviewed
Next slice   Change notification list UI (requires approval)
Decision     Continue | Revise | Stop
Enter fullscreen mode Exit fullscreen mode

The card points to a specific change rather than asking you to trust an agent's self-assessment. A test of preference persistence does not establish that the control behaves correctly in a browser. Even a passing browser interaction wouldn't settle keyboard or screen-reader behavior by itself.

If the agent changes the toggle after that check, the runner must reassess which evidence still applies. Otherwise the UI has turned a passing result for an older candidate into a passing result for whatever code happens to be on screen now.

Make the three buttons mean different things

Continue accepts the reviewed slice and authorizes only the stated next slice. If the runner is free to change unrelated files or launch a deployment, the UI shouldn't call this a bounded approval.

Revise changes the target before more work begins. In the example, the operator might ask for the missing browser scenario, or request a focused fix to the toggle without touching the list. The resulting candidate needs its own review; editing the work does not magically preserve every old check.

Stop prevents new work from being scheduled. It should also show what has already been issued to tools and what remains uncertain. Stopping a model response cannot be assumed to cancel an in-flight command, undo a file write, or reverse a deployment. A product that offers Stop needs an executor-side boundary and a way to reconcile effects that crossed it.

An attractive frontend can lie by accident here. If Continue, Revise, and Stop are labels above one unconstrained agent loop, none of those buttons gives the operator control.

Why this isn't another final approval screen

A product-refactor account on DEV Community describes two principal agents, three review milestones, and checks that included tests and browser workflows. GitHub's account of its Copilot runtime migration describes shipping a much larger Rust change incrementally and treating end-to-end tests as critical. Neither account establishes a universal checkpoint UI or a productivity benchmark. The practical connection is narrower: an increment gives reviewers something they can accept or send back before the next change widens the scope.

Meanwhile, Foundry Toolkit's release notes describe follow-up and stop controls for active hosted-agent responses, plus inspection updates. That is evidence that runtime steering is getting product attention, not proof that pressing Stop gives atomic cancellation or that Foundry implements the checkpoint proposed here. A Hacker News discussion of coding-agent workflows also shows disagreement about letting agents handle correctness-sensitive tasks. A checkpoint doesn't resolve that disagreement; it makes the acceptance decision explicit.

For frontend builders, the presentation still matters. If the checkpoint itself is a generated interface, compare component and renderer approaches in Generative UI resources. Then keep the actual approval action tied to trusted application state and runner enforcement, not a model-written sentence in a card. The resource is for UI patterns, not evidence that any stop control is safe.

Keep the gate proportionate

A one-line copy edit probably doesn't need a modal. A change that affects saved preferences, browser behavior, or a user's next action deserves a pause. Pick a point where a reviewer can still reject the current slice without first untangling several more edits.

Before pressing Continue, I want to know which candidate I'm accepting and which behavior was checked. If the evidence isn't there, Stop must keep the next slice from starting. Otherwise the agent is moving faster than its review loop.


Source notes

Top comments (0)