DEV Community

Cover image for The PAOVR Loop: The Real Agent Loop That Actually Finishes Jobs
Eduard
Eduard

Posted on Originally published at auditme.dev

The PAOVR Loop: The Real Agent Loop That Actually Finishes Jobs

Plan → Act → Observe → Verify → Repair

Stop building agents that narrate completion. Start building systems that prove it.

In 2026, the conversation has moved beyond “prompt engineering is dead.” What matters now is quieter, harder, and far more useful: loop engineering.

Many agents are already competent at individual steps. They still fail on longer jobs because the surrounding system has no enforceable definition of “done.”

We have all seen the pattern: an agent burns tokens, makes a plausible tool call, produces a confident summary, and declares success — while the actual requirement was only partially satisfied, the evidence was never checked, or a failed path quietly changed the goal.

The model is not always the problem. Often, the missing piece is the contract around it.

This is a production-grade field guide to the PAOVR Loop — Plan → Act → Observe → Verify → Repair — a practical control pattern designed to make long-horizon agent work more reliable.

It builds on ideas from ReAct, Plan-and-Solve, modern agent harnesses, and the hard lessons of teams running agents in production rather than demos.

You will leave with:

  • A precise anatomy of the PAOVR Loop for long-horizon tasks
  • The actual prompts we use in production
  • JSON contracts and TypeScript interfaces that agents and runtimes can both consume
  • Real 2026 implementation stacks (Next.js, TypeScript, Supabase, Vercel)
  • Vector memory patterns that turn amnesiac agents into compounding workers
  • Failure patterns that still dominate and how to kill them
  • A one-week install plan you can run on your own stack

This is not theory. It is the difference between an agent that talks about finishing and one that proves it finished.

Table of Contents

  1. Why Most Agents Still Fail in 2026
  2. The Shift from Prompting to Loop Engineering
  3. The PAOVR Loop: Plan → Act → Observe → Verify → Repair
  4. Stage 1 — Plan: Stop Asking Agents to Think. Ask Them to Graph
  5. Stage 2 — Act: Atomic Execution with Tool Contracts
  6. Stage 3 — Observe: Grounding in Reality
  7. Stage 4 — Verify: The Step Almost Everyone Skips
  8. Stage 5 — Repair: Recovery Without Restarting from Zero
  9. JSON Contracts That Survive Production
  10. The Prompts We Actually Use in Production
  11. Context Engineering Inside the Loop
  12. Circuit Breakers, Budgets, and Stopping Conditions
  13. Failure Patterns I Keep Seeing in 2026
  14. How Real Production Systems Use This Loop
  15. A One-Week Install Plan
  16. Ship Checklist
  17. What to Do in the Next 15 Minutes
  18. Further Reading, People & Tools
  19. Frequently Asked Questions

1. Why Most Agents Still Fail in 2026

The failure mode has shifted.

In 2023–2024, models were often simply wrong. In 2026, many models are much better at atomic tasks. The system can still fail because there is no enforceable definition of “done.”

Typical symptoms:

  • The agent produces a beautiful plan and then freestyles the execution.
  • It marks a task complete because the last tool call returned something.
  • It never re-checks the original success criteria after the final action.
  • Context grows until the original goal is buried under tool noise.
  • When something breaks, the agent rewrites the entire plan instead of repairing the broken leaf.
  • A repair changes the acceptance criteria without making that change explicit.

The common root cause is the same: the loop has no Verify stage with teeth, and no immutable completion contract protecting the original goal.

ReAct (Yao et al., 2022) taught us to interleave Thought → Action → Observation. That was necessary. It was not sufficient for long-horizon production work. Plan-and-Solve (Wang et al., 2023) added an explicit planning phase. Modern agent harnesses add budgets, tools, worktrees, and stopping conditions.

The missing discipline is making verification and repair part of the control system rather than optional model behavior.

If your agent cannot answer the question “How do I know this is finished?” with evidence instead of narration, it is not finished.

2. The Shift from Prompting to Loop Engineering

Prompt engineering optimized the single turn.

Loop engineering optimizes the entire trajectory.

The people building reliable agents increasingly talk about:

  • Boris Cherny (Claude Code, Anthropic): “I don’t prompt Claude anymore. I have loops running that prompt Claude.”
  • Addy Osmani and the broader community: loop engineering as a discipline.
  • Anthropic’s guidance on the Agent SDK / Claude Code: gather context → take action → verify work → repeat.
  • Production systems: circuit breakers, max turns, cost thresholds, external verifiers, and explicit stopping conditions.

The unit of design is no longer “the perfect system prompt.”

It is the control loop that keeps the model inside a contract until the contract is satisfied, evidence is insufficient, or the budget is exhausted.

This article is about that loop — specifically the PAOVR version of it.

3. The PAOVR Loop: Plan → Act → Observe → Verify → Repair

Here is the minimal reliable shape:

PLAN
  ↓
ACT (one atomic step)
  ↓
OBSERVE (real tool / environment feedback)
  ↓
VERIFY (against explicit done_when)
  ↓
  ├─ satisfied → next task or finish
  └─ not satisfied → REPAIR → back to ACT or re-plan only the affected subtree
Enter fullscreen mode Exit fullscreen mode

This is the PAOVR Loop.

Key design rules:

  1. One atomic action per Act. Prefer 1–3 tool calls maximum.
  2. Every task has a crisp done_when. If you cannot write it, the task is not ready.
  3. Verify is external or at least independent. The same model that generated the work should not be the only judge.
  4. Repair is local. Do not throw away the entire plan because one leaf failed.
  5. Hard stopping conditions always exist. Iteration limit, cost limit, repeated identical failure, context budget.
  6. The completion contract is immutable by default. Repair can change the execution path, but it cannot silently change what counts as success.

Outcome and execution health are separate

A task has two independent dimensions:

Outcome: Did the task satisfy its original completion contract?

Execution health: How did we get there?

A repaired task can therefore be successful while still having a degraded execution path.

For example:

outcome = satisfied
execution_health = repaired
Enter fullscreen mode Exit fullscreen mode

This means the original requirement was met, but the system had to recover from a failed path.

That is different from:

outcome = satisfied
execution_health = clean
Enter fullscreen mode Exit fullscreen mode

which means the task satisfied its contract without requiring repair.

Other execution-health states can include:

clean
repaired
degraded
budget_exhausted
authorization_blocked
Enter fullscreen mode Exit fullscreen mode

“Done” answers whether the outcome was achieved. Execution health answers whether the path remained healthy. Never collapse these into a single boolean status.

4. Stage 1 — Plan: Stop Asking Agents to Think. Ask Them to Graph

Planning is no longer “think step by step.”

It is the production of an executable graph.

What a good plan looks like

  • Goal stated as an observable outcome
  • Explicit assumptions
  • Clarifying questions only when the cost of being wrong is high
  • Tasks that are leaf-level (doable in 1–3 tool calls)
  • Dependencies declared
  • Every task has a done_when string that a later verifier can check
  • Risks listed
  • The completion contract is defined before execution begins

Planner prompt we actually use

Act as the Task Planner. You do not execute. You only produce an executable plan.

Rules:
1. Split the goal into atomic steps.
2. One step = one action or one tightly related group of tool calls (max 3).
3. Declare dependencies with task IDs.
4. Every step must have a crisp done_when that can be verified later.
5. If critical information is missing, list assumptions and clarifying_questions. Do not invent facts.
6. Output strict JSON only. No prose essay.

Return exactly this schema:
{
  "goal": "string",
  "assumptions": ["string"],
  "clarifying_questions": ["string"],
  "tasks": [
    {
      "id": "t1",
      "title": "string",
      "description": "string",
      "depends_on": ["t0"],
      "tool_hint": "none|search|code|browser|api|file",
      "done_when": "observable condition that proves completion"
    }
  ],
  "risks": ["string"]
}
Enter fullscreen mode Exit fullscreen mode

This planner is deliberately dumb about execution. That is the point. Separation of concerns is what keeps the system debuggable.

5. Stage 2 — Act: Atomic Execution with Tool Contracts

The Executor receives one task, the current plan state, and any previous observations. It is forbidden from jumping ahead.

Executor prompt

Act as the Executor Agent.
Take exactly one next task from the plan. Do not jump ahead. Do not invent missing data.

Inputs you will receive:
- plan JSON
- current_task_id
- previous tool results / observations (if any)

Method:
1. Re-read the done_when for the current task.
2. If you are blocked on missing data, request the cheapest tool or mark status blocked.
3. Perform the smallest useful action that moves the task forward.
4. Return structured output only:

## Action
(what you did)

## Evidence
(raw tool output or observation — never paraphrase away the truth)

## Status
done | partial | blocked

## Residual risks
(any new risks introduced)

## Next recommendation
(only if status is not done)
Enter fullscreen mode Exit fullscreen mode

The Executor never decides the overall goal is finished. That decision belongs to the outer loop after verification.

6. Stage 3 — Observe: Grounding in Reality

Observation is the only place the model is allowed to see the real world.

Rules that still matter in 2026:

  • Never let the model invent tool output. The runtime supplies it.
  • Prefer structured tool responses over free text when possible.
  • Keep the observation window small and high-signal. Context rot is real.
  • Log every observation with a timestamp and tool name. You will need it for debugging.
  • Preserve raw evidence separately from model-generated summaries.

This is the stage that turns ReAct from a clever prompt into a more reliable control system.

7. Stage 4 — Verify: The Step Almost Everyone Skips

Verification is the difference between an agent that claims success and one that demonstrates it.

What “done” actually means

A task is done only when its done_when is true and the evidence supports that claim.

The verifier should preferably be:

  • A separate model call with a different system prompt, or
  • An external checker (tests, linter, schema validator, SEO score, human review), or
  • A deterministic function when the domain allows it.

Verifier prompt

Act as the Verifier. You do not generate new work. You only judge whether the current task is complete.

You receive:
- original task (including done_when)
- action taken
- evidence / observation
- any claimed result

Rules:
1. Quote the done_when.
2. Decide: satisfied | not_satisfied | insufficient_evidence.
3. If not_satisfied, name the single cheapest next check or repair.
4. Never accept narration as proof. Require evidence.
5. Never rewrite or weaken the original done_when.
6. Output strict JSON:

{
  "task_id": "...",
  "done_when": "...",
  "verdict": "satisfied|not_satisfied|insufficient_evidence",
  "evidence_summary": "one or two sentences",
  "missing": ["what is still required"],
  "recommended_repair": "smallest next action or null"
}
Enter fullscreen mode Exit fullscreen mode

This is the stage that prevents the polite lie.

8. Stage 5 — Repair: Recovery Without Restarting from Zero

When Verify returns not_satisfied or insufficient_evidence, the system has two clean options:

  1. Local repair — re-run or adjust only the failed leaf.
  2. Subtree re-plan — only when dependencies or assumptions inside that subtree have actually changed.

Never throw away the entire plan because one step failed. That is how agents waste tokens and lose trust.

The completion contract is immutable

A repair may change how the work is performed.

It may not silently change what counts as success.

The original done_when is the acceptance contract for the task and its repair descendants.

When a subtree is re-planned:

  1. Preserve the original task ID and completion contract.
  2. Create a new plan version for the replacement subtree.
  3. Link replacement tasks to the original task.
  4. Carry the original done_when forward unchanged.
  5. Record every changed assumption or dependency.
  6. Preserve evidence from the failed subtree rather than deleting it.
  7. Require explicit approval before weakening or changing the completion contract.

The core rule is:

Re-planning changes HOW the work is performed. It does not change WHAT counts as success.

Repair example

Suppose the original task is:

Task: Deploy API
done_when:
Production deployment is live, all tests pass, and p95 latency is below 200 ms.
Enter fullscreen mode Exit fullscreen mode

The deployment fails because a dependency is unavailable.

A repair can replace the execution path:

t3a → provision alternate dependency
t3b → deploy API
t3c → run production verification
Enter fullscreen mode Exit fullscreen mode

But t3c must still verify the original contract:

Production deployment is live,
all tests pass,
and p95 latency is below 200 ms.
Enter fullscreen mode Exit fullscreen mode

The repair is allowed to change the path. It is not allowed to quietly turn that into:

API deployment command completed.
Enter fullscreen mode Exit fullscreen mode

That distinction is what prevents a repair planner from converting failure into a false success.

Repair record

For production systems, keep an explicit repair record:

interface RepairRecord {
  repair_id: string;
  parent_task_id: string;
  original_contract: string;
  original_contract_version: number;
  new_plan_version: number;
  changed_assumptions: string[];
  reason: string;
  evidence: string[];
}
Enter fullscreen mode Exit fullscreen mode

The failed path becomes part of the audit trail rather than disappearing from history.

9. JSON Contracts That Survive Production

Free-form text is fine for humans. Agents need schemas.

Here is a minimal production-oriented run schema:

{
  "run_id": "uuid",
  "goal": "...",
  "outcome": "in_progress|satisfied|not_satisfied|insufficient_evidence",
  "execution_health": "clean|repaired|degraded|budget_exhausted|authorization_blocked",
  "tasks": [
    {
      "id": "t3",
      "status": "pending|running|done|partial|blocked|failed",
      "verdict": "pending|satisfied|not_satisfied|insufficient_evidence",
      "attempts": 2,
      "contract_version": 1,
      "plan_version": 2,
      "parent_task_id": "t3",
      "last_evidence": "...",
      "verified_at": "ISO timestamp"
    }
  ],
  "cost_so_far": {
    "tokens": 12840,
    "usd_estimate": 0.41
  },
  "circuit_breaker": {
    "max_turns": 40,
    "max_cost_usd": 5.0,
    "identical_failure_limit": 3
  }
}
Enter fullscreen mode Exit fullscreen mode

The important change is that outcome and execution health are independent.

A run can be:

outcome = satisfied
execution_health = repaired
Enter fullscreen mode Exit fullscreen mode

without pretending it was a clean execution.

TypeScript contracts

type TaskVerdict =
  | "pending"
  | "satisfied"
  | "not_satisfied"
  | "insufficient_evidence";

type ExecutionHealth =
  | "clean"
  | "repaired"
  | "degraded"
  | "budget_exhausted"
  | "authorization_blocked";

interface Task {
  id: string;
  title: string;
  description: string;
  depends_on: string[];
  tool_hint: "none" | "search" | "code" | "browser" | "api" | "file";

  // Immutable acceptance contract
  done_when: string;
  contract_version: number;

  // Execution state
  status?: "pending" | "running" | "done" | "partial" | "blocked" | "failed";
  verdict?: TaskVerdict;
  execution_health?: ExecutionHealth;

  attempts?: number;
  last_evidence?: string;
  verified_at?: string;

  // Repair / re-plan lineage
  parent_task_id?: string;
  plan_version: number;
}

interface AgentPlan {
  goal: string;
  assumptions: string[];
  clarifying_questions: string[];
  tasks: Task[];
  risks: string[];
}

interface RunState {
  run_id: string;
  goal: string;

  outcome:
    | "in_progress"
    | "satisfied"
    | "not_satisfied"
    | "insufficient_evidence";

  execution_health: ExecutionHealth;

  tasks: Task[];

  cost_so_far: {
    tokens: number;
    usd_estimate: number;
  };

  circuit_breaker: {
    max_turns: number;
    max_cost_usd: number;
    identical_failure_limit: number;
  };
}
Enter fullscreen mode Exit fullscreen mode

These interfaces become the source of truth between your orchestrator, execution layer, and logging system.

One implementation detail matters: TypeScript interfaces alone do not enforce runtime JSON from an LLM. In production, validate model output at runtime with a schema validator such as Zod or JSON Schema before mutating state.

10. The Prompts We Actually Use in Production

You already have the three core ones (Planner, Executor, Verifier).

Here is the outer loop controller that ties them together:

Act as the Loop Controller. You own the overall trajectory.

Your only job:
1. Load or create the plan.
2. Select the next ready task (dependencies satisfied, status not done).
3. Hand it to Executor.
4. Feed the result to Verifier.
5. On satisfied → record the verified outcome and continue.
6. Record execution health separately from the outcome.
7. On not_satisfied or insufficient_evidence → trigger Repair (local first).
8. Enforce circuit breakers before every new turn.
9. Never allow Repair or re-planning to weaken the original completion contract.
10. When all tasks satisfy their original contracts, emit final result + execution health + residual risks.
11. Never invent completion.

You speak only in structured status updates and JSON state.
Enter fullscreen mode Exit fullscreen mode

These four prompts form a deployable skeleton for the PAOVR Loop. Keep them versioned in git the same way you version any other critical configuration.

11. Context Engineering Inside the Loop

Context is a finite resource. In long runs it becomes a primary failure mode.

Practical rules that still hold:

  • Keep the master policy (role, constraints, output contract) stable and cached.
  • Give the Executor only the current task + recent observations + the original done_when.
  • Summarize or offload completed tasks instead of replaying the entire history.
  • Prefer fresh context for pure execution workers and accumulated context only for the planner/orchestrator.
  • Measure context fill. When it crosses ~60–70% of the useful window, force a compression or checkpoint step.

This is why many robust agent systems treat the file system, git, and external memory as first-class context tools rather than dumping everything into the prompt.

Vector Memory as a First-Class Citizen

Context windows are large, but dumping everything into them destroys attention. External memory architectures can reduce that pressure.

A practical pattern:

  1. After every completed (or failed) task, embed a short structured summary of what happened, the evidence, and the outcome.
  2. Store those embeddings in a vector store. pgvector is a practical choice when vector search can live next to relational state.
  3. Before the Plan stage of a new run, the orchestrator performs a micro-RAG retrieval against the agent’s own historical executions.
  4. The retrieved constraints are injected into the Planner’s context as hard lessons (“previous attempts failed when the shadow-DOM selector timed out; prefer the data-testid path”).

The effect can compound. An agent that repeatedly failed to interact with a particular UI element across past sessions no longer has to rediscover the failure mode.

Pair embeddings with a fast, high-quality text-embedding model. Gemini embedding models are one option for balancing cost and quality. Keep the retrieval budget tiny — usually the top 3–5 most relevant past failures or successes are enough. Anything more can re-introduce context rot under a different name.

12. Circuit Breakers, Budgets, and Stopping Conditions

A loop without hard stops is a liability.

Minimum set:

Signal Typical setting Enforcement
Max turns / iterations 20–60 depending on task Runtime
Max cost (USD or tokens) Task-specific budget Runtime
Identical failure streak 2–3 Instruction + runtime
Context budget 70% of useful window Instruction
Wall-clock timeout Optional Runtime

When a breaker trips, the agent must:

  1. Stop new actions.
  2. Return partial results that were already verified.
  3. State clearly what triggered the stop and what remains open.
  4. Escalate if a human gate exists.

Partial verified work is always more valuable than a confident hallucination.

13. Failure Patterns I Keep Seeing in 2026

  1. Narrated completion — the model says “done” without evidence.

    Fix: hard Verify stage with external or independent judgment.

  2. Plan that is actually a novel — tasks that still require a short essay of instructions.

    Fix: keep splitting until each leaf is 1–3 tool calls.

  3. Context rot — original goal buried under 30 tool observations.

    Fix: aggressive pruning + separate orchestrator context + vector memory for long-term lessons.

  4. Repair by total rewrite — one failure causes the agent to discard everything.

    Fix: local-first repair and explicit subtree lineage.

  5. Contract weakening during repair — the replacement plan quietly lowers the acceptance bar.

    Fix: make done_when immutable by default, version the contract, preserve the original evidence, and require explicit approval for contract changes.

  6. Missing done_when — “make it good” or “optimize the page.”

    Fix: refuse to accept a task without an observable completion condition.

  7. Tool hallucination — model invents tool results.

    Fix: runtime always supplies Observation; model is never allowed to generate it.

  8. Infinite polite loops — agent keeps “trying one more thing.”

    Fix: circuit breakers with identical-failure detection.

These are recurring production failure modes. The exact frequency varies by system, model, tools, and task design.

14. How Real Production Systems Use This Loop

The pattern appears, under different names, in systems that ship:

  • Claude Code and the Anthropic Agent SDK / Claude Code docs — gather → act → verify → repeat, with explicit loop and stopping concepts.
  • Coding agents that treat the test suite as the verifier.
  • Research agents that force a verification step against sources before claiming a fact.
  • Content and SEO pipelines that run a quality gate after generation.

The 2026 Implementation Layer

The theory maps directly to modern stacks. You do not need a massive monolithic Python backend to run this loop cleanly.

A common high-leverage architecture in 2026:

  • Orchestration: Next.js App Router (or a lightweight server layer) owns the Loop Controller. Strict TypeScript contracts define the state model.
  • State & Logs: Supabase (Postgres + pgvector) stores run state, task history, and vector memory of past executions.
  • Execution: Serverless / edge functions on Vercel handle individual Act steps. This keeps the execution surface small and isolates failures.
  • Tool surface: Many teams standardize on the Model Context Protocol (MCP) so agents can talk to tools through a consistent interface.
  • Local development & coding agents: The same philosophy applies to tools like OpenCode and Cline. They write code, observe terminal output, verify against tests or linters, and repair locally rather than blindly restarting.

The key insight is that the PAOVR Loop is language-agnostic. Once you have typed contracts, runtime validation, and a reliable state store, the same shape can work for coding agents, research agents, and domain-specific crawlers.

Case Study: Surviving the Chaos of Web Crawling

Let’s look at a real production environment. When building the crawler pipeline for the AI SEO platform AuditMe, the biggest nightmare wasn’t parsing HTML — it was the unpredictability of the web. Sites time out, DOMs shift, JavaScript-heavy pages render differently, and standard linear scripts can break under noisy conditions.

To address this, the auditing engine was structured around PAOVR-style control. Instead of a monolithic script, the system uses Next.js App Router and Supabase to orchestrate atomic tasks. For example, if you run a URL through the free Website SEO Checker, the workflow can trigger multiple stages of planning, execution, observation, verification, and repair.

If a check fails (for example, an API timeout during a heavy DOM render), the system can catch the failure in verification, trigger a provider failover or local repair, and continue. Only the affected leaf needs to be retried.

The broader lesson is not that every crawler needs the same implementation. It is that unreliable external environments require explicit evidence, bounded retries, and local recovery.

15. A One-Week Install Plan

Day 1

Write the three core prompts (Planner, Executor, Verifier). Run them manually on a simple multi-step task. Measure where the model tries to skip Verify.

Day 2

Add strict JSON schemas (or TypeScript interfaces plus runtime validation) and a simple state object. Make the outer loop refuse to continue without valid state.

Day 3

Introduce one external verifier (tests, schema check, or a second model call). Force the system to use it.

Day 4

Add circuit breakers: max turns, cost, identical failure. Test them by deliberately breaking a tool.

Day 5

Implement local-first Repair. Confirm that a single failed leaf does not destroy the whole plan. Add explicit task lineage and contract versioning. Optionally wire a minimal pgvector memory store for past failures.

Day 6

Run a real 20–40 step task. Log every observation and verification. Identify the highest-friction stage.

Day 7

Write the one-page internal playbook for your team. Version the prompts and schemas. Put the state schema in git.

By the end of the week you should have a PAOVR Loop that is bounded, observable, and materially easier to trust than an agent that relies on narration alone.

16. Ship Checklist

Before you call any agent “production”:

  • [ ] Every task has an explicit done_when
  • [ ] Planner and Executor are separate
  • [ ] Verify stage exists and is independent
  • [ ] Repair is local-first
  • [ ] Original completion contract is immutable across repairs
  • [ ] Replacement subtrees preserve contract lineage
  • [ ] Contract changes require explicit authorization
  • [ ] Repair history and failed-path evidence are preserved
  • [ ] Outcome and execution health are tracked separately
  • [ ] Circuit breakers are enforced by the runtime, not just the prompt
  • [ ] Observations are never invented by the model
  • [ ] Completed work is preserved and evidenced
  • [ ] Cost and turn budgets are visible
  • [ ] Partial results are returned on early stop
  • [ ] Prompts and schemas are versioned
  • [ ] Long-term lessons are stored outside the context window (vector memory or equivalent)
  • [ ] LLM-generated JSON is validated at runtime before state mutation

If any critical box is unchecked, the agent is still closer to a demo than a production system.

17. What to Do in the Next 15 Minutes

  1. Copy the Planner prompt into your current agent stack.
  2. Take one real task you care about and force it to emit the JSON plan schema.
  3. Write a done_when for the first three leaf tasks that a stranger could verify.
  4. Add a single Verify call after the first Act.
  5. Add a simple immutable contract field and task lineage before testing Repair.
  6. Run it once and look at the difference between narration and evidence.

That is the difference between “it usually works” and “I can trust it when I’m not watching.”

18. Further Reading, People & Tools

Foundational papers

People & practice

  • Boris Cherny (Claude Code, Anthropic) — loop-first mindset
  • Addy Osmani — loop engineering discussions and practice
  • Anthropic Agent SDK / Claude Code docs
  • OpenAI’s agent and Codex documentation

Tools & platforms worth watching

19. Frequently Asked Questions

Isn’t this just ReAct with extra steps?

ReAct is the necessary interleaving of reasoning and acting. The PAOVR Loop adds explicit planning with contracts, independent verification, and controlled repair. Those additions make the control loop better suited to long-horizon work.

Do I still need a strong system prompt?

Yes. The prompts above are the system prompts. They are focused on policy and contracts rather than personality.

Can the same model do Plan, Act, and Verify?

It can, but independence is preferable. Even when the same base model is used, separate roles, prompts, and ideally different verification mechanisms reduce the chance that the system simply approves its own work.

What about multi-agent systems?

The same loop still applies. The orchestrator runs the outer PAOVR Loop; specialist agents become the Act stage for particular tools or domains.

How do I add long-term memory without exploding context?

Use vector memory (pgvector + embeddings) and retrieve only the top few relevant past failures or successes before planning. Keep the retrieval budget tiny.

How do I prevent Repair from weakening the goal?

Treat the original done_when as an immutable completion contract. Give replacement tasks explicit lineage to the original task, carry the contract forward unchanged, preserve failed-path evidence, and require an explicit authorization step before changing acceptance criteria.

How do I know when to stop adding stages?

When the agent can finish a meaningful multi-step task, survive a tool failure, preserve verified partial results, and stop under a hard budget — stop. Further complexity should earn its place by reducing a measurable failure mode.

Final note

The agents that will still be running in production in 2027 are unlikely to be the ones with the cleverest personality block. They will be the ones whose loops enforce a contract, demand evidence, remember past failures, and know how to repair without starting over.

Build the PAOVR Loop.

Version the contracts.

Verify everything.

Give the agent a memory that compounds.

Then the model can finally do what we have been asking it to do for years: finish the job.

Written for practitioners who ship. Updated for the 2026 agent landscape.

Top comments (5)

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@edo911, the Verify gate is the strongest part of PAOVR, especially requiring an observable done_when before execution. I’d keep task outcome separate from path health: a repair can eventually satisfy done_when while still exceeding retry, cost, or authorization budgets, and those runs should not collapse into the same “done” state as a clean execution. What evidence do you preserve across a subtree re-plan so the replacement plan cannot silently weaken the original completion contract?

Collapse
 
edo911 profile image
Eduard •

Great catch, Raju. You’re right that the original version didn’t make this distinction explicit.

I’ve updated the article to separate the task outcome from execution health. A task can now be satisfied while still being marked as repaired or degraded.

I also added an immutable completion contract: subtree re-planning can change how the work is executed, but it cannot silently weaken the original done_when. The original contract, evidence, task lineage, and repair history are preserved, and changing the acceptance criteria requires an explicit decision.

That was an important production detail missing from the first version. Thanks for pointing it out.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.