A coding agent can run tests, read the failure, edit code, and run the tests again. That sounds like a simple loop. Without a control structure, though, the same loop can become a string of untracked edits, repeated commands, and increasingly confident summaries.
A dependable workflow treats every action as a hypothesis about the system and every result as evidence that may change the next step. It also names the points where an agent must stop. The goal is not to keep the loop running until it says “done”; the goal is to produce a useful state transition with evidence a reviewer can inspect.
Model the work as states
An engineering episode can be represented as a small state machine:
| State | Question to answer | Exit condition |
|---|---|---|
| Intake | What outcome, scope, and constraints were requested? | The task contract is understood; material ambiguity is recorded |
| Observe | What do the repository, tests, logs, and docs actually show? | Relevant facts and unknowns are separated |
| Model | What failure or change mechanism best explains the evidence? | There is a testable working hypothesis |
| Plan | What bounded actions can test that hypothesis? | Actions fit the authority and scope |
| Act | What changed, and what commands or tools ran? | The planned action produced a result or a stop condition |
| Evaluate | Does the result support the acceptance conditions? | Evidence is accepted, contradicted, stale, or insufficient |
| Exit | Is the episode complete, blocked, or ready for human judgment? | A truthful status and evidence report are returned |
The names are not sacred. A team can combine or rename states. What matters is that observation is not confused with inference, an attempted action is not confused with success, and a model’s summary is not substituted for a tool result.
Keep facts, hypotheses, and decisions apart
Suppose a test fails after a dependency update. A disciplined trace might say:
- Observed: the token-refresh test fails with an unexpected 401.
- Hypothesis: the new client no longer retries after refreshing credentials.
- Next action: inspect the new retry behavior and add a focused regression test.
- Decision boundary: if preserving behavior requires changing the public authentication contract, stop for a human decision.
The separation is useful because a plausible explanation can be wrong. If the agent records its hypothesis as though the logs established it, later actions inherit a false premise.
When multiple explanations remain possible, design the smallest useful experiment. Read the relevant library documentation, isolate the test, compare the old and new behavior, or inspect a sanitized trace. Avoid changing several unrelated variables at once; otherwise the outcome may not tell you which change mattered.
Make evidence expire when its subject changes
Verification is tied to the thing that was checked. If the code changes after a test run, that test result no longer describes the current code. If a test command is rerun with a different configuration, the old and new results are not interchangeable. If an agent changes the acceptance test itself, a green result needs a separate review of that change.
Treat an evidence item as a record with at least:
| Evidence field | Example |
|---|---|
| Subject | Working-tree revision or built artifact |
| Check | Exact command, test suite, or policy |
| Context | Runtime, fixture, configuration, and relevant dependencies |
| Result | Exit status plus failures, skips, or unknowns |
| Time/order | Whether it ran before or after the last relevant change |
An important rule follows: after a material edit, rerun the checks that the edit could invalidate. A test from before the patch can explain the starting condition; it cannot qualify the final patch.
This does not require rerunning every expensive job after every keystroke. Match the check to the risk and the changed surface. A documentation-only change might need a link check and rendered preview. A serialization change may need compatibility fixtures and migration tests. The completion report should say what was not rerun and why.
Use retries to learn something
Retries are useful when the next attempt differs in a way that could address the observed failure. Repeating an identical command against an unchanged state usually gives the same evidence and consumes attention.
Before a retry, ask:
- What did the previous attempt establish?
- What will be different in this attempt?
- What result would change the working hypothesis?
- How many attempts or how much runtime did the contract allow?
If a command fails because the local service is not running, starting that service is a meaningful next step. If it fails with the same assertion after two unmodified runs, another identical run is not a plan. If a tool call may have partially changed an external system, first determine whether repeating it is safe; a non-idempotent action can create duplicate tickets, releases, or payments.
Bound retries by both count and consequence. For example: retry a deterministic local check once after a relevant fix; stop after a repeated infrastructure failure; require a person before repeating any external action whose first outcome is unknown.
Put the stop decision in the workflow
Stop and escalate when:
- the task requires a product or policy choice the contract did not delegate;
- the available evidence conflicts or cannot identify the tested state;
- a required credential or production data would be needed;
- the same failure persists without a new diagnostic signal;
- a proposed fix expands the scope or weakens an invariant;
- an external side effect may have happened but cannot be confirmed.
These are workflow outcomes, not model moods. “I am confident” should not override a failed check or missing permission. Likewise, low confidence alone need not block a reversible, low-risk investigation if the agent can gather better evidence within its scope.
Capture enough of the trajectory to review
A useful trace records the task identifier, relevant inputs, tool calls, changed paths, check results, handoffs, and final status. Keep secrets and unnecessary personal data out of that record. The purpose is not to preserve every token; it is to let someone answer what the agent saw, what it did, what it learned, and why it stopped.
For teams using an agent platform, inspect what its trace actually captures. OpenAI’s current evaluation guide describes traces as end-to-end records of model calls, tool calls, guardrails, and handoffs. That is a useful example of a trace surface, but no vendor trace by itself proves that the application’s acceptance criteria were valid or that a consequential action was authorized.
A practical completion report
At the end of the episode, return a short, structured report:
- Status: complete, incomplete, blocked, or escalated.
- Change: what files or external records changed.
- Evidence: exact checks and results, including failures and skips.
- Assumptions: choices made because the task left a gap.
- Remaining risk: behavior not exercised or not covered by the evidence.
- Next decision: the specific question a person needs to answer, if any.
This format makes a failed episode useful too. “Blocked because the provider’s migration guide leaves token refresh behavior unspecified” is actionable. “Couldn’t finish” is not.
The control-loop test
Before adopting a workflow, walk through one ordinary task and one deliberately awkward case. Ask whether it can distinguish a fact from a hypothesis, invalidate old evidence, limit repeated actions, and reach a truthful stop state. If its only terminal condition is a success message, the loop is incomplete.
An agentic engineering loop is effective when feedback changes what happens next and authority constrains what may happen at all. That is how repeated tool use becomes a controlled engineering process instead of a long conversation with a terminal.
References
- Hassan et al., “A Roadmap for Agentic Software Engineering”
- OpenAI API documentation, “Evaluate agent workflows”
Source, license, and AI assistance
This article is based on the agentic engineering loop and evidence concepts in Part II of From Vibe Coding to Agentic Software Engineering, whose source record credits ChatGPT as preparer and identifies CC BY-NC-SA 4.0. This version is substantially reorganized and expanded with original examples and a practical workflow, and is shared under the same license: CC BY-NC-SA 4.0.
AI disclosure: The article text was generated primarily by AI. A human publisher supplied the topic, source material, and editorial direction, and remains responsible for checking claims and examples before publication.
Top comments (0)