Verdict: Feedback loops in agentic AI systems are what turn a single model call into an agent, and the evaluator-optimizer loop (a generator model judged by an evaluator with access to tests, schemas, or live data) wins for any task where correctness can be checked. The unchecked self-critique loop, where the same model re-reads its own answer with no external signal, measurably degrades accuracy on reasoning tasks (Huang et al., ICLR 2024) and should be capped at one pass, reserved for style and format polish. Choose the loop that is grounded in a real-world check; treat self-critique as polish, never as verification.
TL;DR
- Winner: the evaluator-optimizer loop, defined by Anthropic as one model call generating a response while another provides evaluation and feedback in a loop (Anthropic, Building Effective Agents).
- Loser: intrinsic self-critique with no external feedback. Performance on reasoning tasks sometimes degrades after self-correction (arXiv:2310.01798).
- Why earlier results looked better: several self-critique studies used oracle labels to decide when to stop correcting, a signal a production agent does not have (arXiv:2310.01798).
- The fixable middle ground: self-correction trained with multi-turn RL improved MATH by 15.6% and HumanEval by 9.1% over base self-correction (Google, SCoRe).
- Last verified: 2026-09-19.
What role do feedback loops play in agentic AI systems?
A feedback loop is the difference between a prompt and an agent. Anthropic defines an agent as an LLM using tools based on environmental feedback in a loop, gaining ground truth from the environment at each step through tool results and code execution (Building Effective Agents). The loop supplies three things a one-shot call cannot: a signal about whether the last action worked, a place to store what was learned, and a stopping condition.
Reflexion is the clearest research example of the storage part. Instead of updating weights, the agent writes a verbal reflection on task feedback and keeps it in an episodic memory buffer that informs later attempts. On HumanEval it reported 91% pass@1 against 80% for the GPT-4 baseline it was compared with (Shinn et al., arXiv:2303.11366).
How does an evaluator-optimizer loop differ from self-critique?
Both loops generate, judge, and revise. The difference is where the judgement comes from.
| Evaluator-optimizer loop | Unchecked self-critique loop | |
|---|---|---|
| Judge | A separate call, ideally with tool or test access | The same model, same context |
| Signal | External: exit codes, schema errors, diffs, retrieved data | Internal: the model's own confidence |
| Best fit | Tasks where a human could articulate useful feedback | Style, tone, formatting, surface polish |
| Main failure | Cost and latency per iteration | Confident revision of a correct answer into a wrong one |
Anthropic's guidance on when the evaluator-optimizer pattern fits is specific: use it when responses can be demonstrably improved by articulated feedback, and when the model can produce that feedback. The pattern is justified by the existence of a useful critique signal, not by the presence of a second model call.
Why does unchecked self-critique degrade accuracy?
Because the model's error and the model's judgement of that error share a cause. If a chain of reasoning went wrong because the model misread a constraint, re-reading the same context with the same weights tends to reproduce the misreading and then defend it. The ICLR 2024 analysis of intrinsic self-correction, where the model corrects itself based solely on its inherent capabilities without external feedback, found it struggles on reasoning and that performance at times degrades after correction (arXiv:2310.01798).
This does not make self-refinement worthless. Self-Refine, in which the same model generates, critiques, then refines, was preferred over one-shot generation by roughly 20% absolute on average across its task set (Madaan et al., arXiv:2303.17651). The caveat is in the setup: GPT-4 served as both generator and feedback provider, so the gain is strongest on tasks where the critique concerns presentation rather than correctness.
Can training fix self-correction without an external judge?
Partly, and only with training rather than prompting. Google's SCoRe work found prompted self-correction largely ineffective, and instead used multi-turn online reinforcement learning on the model's own generated data, producing a 15.6% gain on MATH with Gemini 1.0 Pro and a 9.1% gain on HumanEval with Gemini 1.5 Flash (arXiv:2409.12917). The practical reading for a team that does not train models: you cannot prompt your way to reliable self-correction, so buy a model trained for it, or supply the external check yourself.
Which model should sit inside the loop?
The cheaper one, more often than teams assume, provided the check is machine-scored. We ran an internal comparison on an identical seven-constraint article-planning task, three trials per model through the Antigravity CLI in headless print mode, with constraints scored programmatically. Gemini 3.8 Flash (High) and Claude Opus 4.6 (Thinking) both scored 17 of 17 on constraint adherence. Median wall time was 23 seconds for Gemini against 67 seconds for Opus (first-party test, n=6, measured 2026-09-19).
That result only holds where a script can grade the output; it says nothing about open-ended judgement. But loop iterations multiply, so a model that is roughly three times faster at equal constraint adherence changes the economics of running four iterations instead of one. If you are picking a coding agent for that seat, our OpenClaude vs Claude Code comparison covers the tradeoffs.
When does a feedback loop cost more than it returns?
When per-step reliability is high enough that iteration adds spend without adding accuracy, or so low that iteration cannot rescue it. A step that succeeds 95% of the time, repeated across ten sequential steps, completes end to end about 60% of the time (reliability multiplication). Loops help by catching failures at each step, but each retry is another billable call.
Three guardrails before the loop ships:
- A hard iteration ceiling. Anthropic recommends explicit stopping conditions such as a maximum number of iterations; two or three is usually enough.
- A machine-checkable exit test. If the evaluator cannot answer "did this pass" with a boolean, you have a conversation rather than a loop.
-
Durable state, not in-memory state. In LangGraph,
interrupt()suspends the graph at the exact call point, state is saved through a checkpointer that can be database-backed in production, and execution resumes withCommand(resume=...)against the samethread_id(LangGraph docs). That is how a loop survives a restart or a wait for human review.
Anthropic's broader advice applies here: find the simplest solution possible, and only increase complexity when needed. A deterministic workflow with one validation step beats a free-running agent for most business tasks, a point we develop in agentic AI versus traditional automation.
The verdict, and who each loop is for
Use an evaluator-optimizer loop with an external check as the default wherever correctness is verifiable: code, structured data extraction, SQL, config generation. This is the pattern behind agents that resolve real GitHub issues in SWE-bench Verified from the pull request description alone. Use self-critique only where the judgement is about form rather than fact, capped at one pass, and never let it overwrite a result that already passed an external check.
If you are designing the surrounding system, the loop is one component among several. Our engineering principles for agentic systems covers the rest, and agentic AI versus AI agents clarifies where orchestration sits relative to the individual agent.
FAQ
Q: What role do feedback loops play in agentic AI systems?
A: They supply the environmental signal that makes an agent an agent: tool results and test outcomes tell the model whether its last action worked, memory carries that lesson forward, and a stopping condition ends the loop. Without an external signal, the loop is only re-prompting.
Q: Is self-critique always worse than an evaluator loop?
A: No. Self-refinement helps on presentation-oriented tasks and was preferred over one-shot generation by roughly 20% absolute on average in the Self-Refine study (Madaan et al.). It is unreliable specifically on reasoning, where correction can degrade the answer.
Q: How many loop iterations should I allow?
A: Set an explicit ceiling, commonly two or three, and log how often the cap is hit. Anthropic recommends stopping conditions such as a maximum iteration count, because unbounded autonomy raises cost and risks compounding errors.
Q: Do I need a second model for the evaluator?
A: Not necessarily a different model, but you do need a different signal. A separate call with test access, schema validation, or retrieved data is what makes the critique useful. The same model with the same context and no tools adds little.
Q: Can reinforcement learning fix self-correction?
A: It can improve it materially. Multi-turn online RL on self-generated data produced gains of 15.6% on MATH and 9.1% on HumanEval over base self-correction in Google's SCoRe work (Google), whereas prompting alone was found largely ineffective.
Q: How do I keep a long-running loop from losing state?
A: Persist it. LangGraph's checkpointer saves graph state at the interrupt point and resumes from the same thread identifier, so a loop can pause indefinitely for human review and survive a restart.
Last verified: 2026-09-19. Figures above are linked to their primary sources; the model timing comparison is our own test, described inline with its method and sample size.
Corrections log: No corrections to date. If you find an error, tell us and we will amend the article and note the change here.
This article was produced with AI assistance and reviewed against the primary sources cited above. Learn how we work on how we work.
Top comments (0)