An agent can edit files, call tools, retry work, and then write: “Done.” That sentence may be useful. It is not an acceptance test.
The practical question is smaller and stricter: who is allowed to decide that a task is complete? If the answer is the same worker that changed the workspace, then the system has a completion claim. It does not yet have an independently checked result.
Bounded Loops & Graphs is an Apache-2.0 agent-harness project built around that distinction. A worker attempts the task. A separate gate decides whether the required property holds. Declared limits stop further work. A receipt records what happened. The model can be Claude Code, Codex, Hermes, another local CLI, or no model at all in a keyless reference loop. The control rule stays the same.
The research paper, Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses, presents the formal model, termination result, spend-bound design, and gate-testing apparatus behind the project.
A worker's completion claim and a gate verdict have different authority.
This article explains the mechanism, shows a first run, and draws the limits clearly. It is aimed at people who already know that capable models can work on code, research, operations, and content, but need a better answer to: what stops the work, what proves the result, and what can a reviewer inspect afterwards?
If you want the underlying vocabulary first, read the inner loop, outer loop, and gate guide. This article starts at the point where a harness must turn those ideas into an enforceable completion decision.
The problem with “the agent says it is done”
Most agent workflows start with a sensible instruction: investigate a problem, make the fix, run a check, and report back. The weak point appears at the end. The worker often decides whether to stop and is also the narrator of the outcome.
That creates a control problem even when the model is competent. A worker can misunderstand the requirement, run the wrong test, interpret an incomplete result as success, or modify the thing that is supposed to evaluate it. A useful harness needs a completion path that does not depend on the worker's own declaration.
Bounded Loops separates the roles:
| Component | Job | What it must not decide |
|---|---|---|
| Worker | Attempts the task in its permitted workspace | Whether the task passed |
| Gate | Checks the declared property against the workspace or artifact | How the worker performs the task |
| Controller | Applies the rule, budgets, and terminal status | Whether a worker's prose sounds convincing |
| Receipt | Preserves the run record for later inspection | Whether to rewrite history |
The worker supplies an artifact; the configured gate evaluates it; the controller applies bounds; the receipt records the decision. The local hash chain is not third-party attestation.
In the current implementation, a worker's agent_claimed_done field can be preserved in the ledger, but the control path does not use it to produce DONE. A gate verdict is required. That is an engineering property of the controller, documented in the project's architecture, rather than a request for the model to be more cautious.
The distinction looks modest until an agent gets access to a real repository. A test that passes is a claim about a test run. A schema that validates is a claim about a concrete artifact. A margin policy that holds is a claim about every relevant row. These are all inspectable properties. “I fixed it” is not the same class of evidence.
A bounded loop in one sentence
A bounded loop is a worker, an independent gate the worker cannot write to, and a set of limits declared before the run starts.
The paper's abstract names three intended guarantees under its stated assumptions: the run finishes under a globally bounded repair model; a node cannot reach DONE without a gate verdict recorded in an append-only hash-chained ledger; and a declared ceiling is enforced inside an attempt rather than only between attempts. Those are design and proof claims about the stated model. They do not establish that every arbitrary agent deployment, integration, or future version is safe.
Three terms are enough to orient a first implementation:
A failed gate can trigger another lap only while declared bounds remain. A passing gate produces DONE; an exhausted bound produces HALT.
Gate: what decides the required property
A gate is a check that has a meaningful failure condition. It can be pytest, a command that checks a data file, a JSON Schema validator, a security scanner, a citation checker, or another mechanical evaluator. A model may help produce the work, but it should not be the only authority that declares the work acceptable.
The gate's independence has a practical component: the worker must not be able to edit the policy or checker it is being judged against. It also has a conceptual component: a checker that merely repeats a value supplied by the worker is still self-attestation, even if it lives in a different file.
Bound: what stops the run
A bound is authority declared before work begins. Depending on the loop or graph, that can include maximum attempts, wall-clock time, token or spend limits when the runner can report them, a no-progress window, or a required approval. The point is not to predict every failure. It is to make the system halt when its declared authority has been used.
This is particularly useful when a task is capable of repeatedly consuming tool calls. “Try again until it works” is an instruction with no stopping rule. A bounded loop makes the stopping rule an executable part of the task contract.
Receipt: what a reviewer can inspect
Each run leaves an append-only, hash-chained record. The receipt can contain the gate verdict, attempts, timing, reported usage, artifacts, and terminal decision. That turns a launch claim or internal handoff from “the agent completed it” into “here is the result, the gate that passed, and the record of the run.”
Receipts are evidence for a particular execution. They do not forecast future cost, quality, or behavior. That distinction is important when operating models change or a new task brings a new failure mode.
Run the smallest real example first
The simplest way to understand the design is to run a small loop where the failure condition is obvious. The current repository documents this keyless first run:
pip install bounded-loops
bl loops install bug-fix-red-green
bl run .bounded-loops/loops/bug-fix-red-green --yes
The installed bug-fix-red-green package contains a planted bug, a stub worker, and a real pytest gate. No API key is required. The stub keeps the setup focused on the control rule: the worker can make the change, but only the test suite can turn the run into DONE.
Expected output includes the declared runner and gate, followed by a gate-passed terminal state:
[bounded-loops] About to run loop 'bug-fix-red-green':
runner : stub
gate : pytest -q
✓ [DONE] gate-passed (laps: 1)
Gate verified: the independent acceptance gate passed after 1 lap.
The exact ledger location is emitted by the CLI. Open that ledger after the run. The useful question is not whether the terminal line is attractive; it is whether you can find the gate verdict and understand what it checked.
Then run the intentionally ungated contrast on a clean repository checkout:
./loops/bug-fix-red-green/wreck.sh
That script is designed to demonstrate the failure mode. The stub claims success while the bug remains, and a separate diagnostic test shows the mismatch. It is a controlled illustration, not a statement that every agent will behave that way. Its value is that it makes the missing control visible: if nothing outside the worker evaluates the required property, a success sentence can end the workflow while the work is still wrong.
Explanatory replay of the shipped keyless example: the worker is a stub; the pytest gate is real. Run the source example and inspect its receipt.
A policy example: repair the catalog, not the verifier
The strongest recipes are not necessarily the largest. They make the acceptance rule and the forbidden shortcut easy to see.
Consider the shipped price-margin-floor loop. Its seed catalog contains three SKUs. SKU-200 costs 20.00 and is priced at 21.50, which is a 7.5% margin. The policy floor is 15%, so the minimum acceptable price is 23.00. The gate reads the catalog and policy and rejects a price below the stated floor.
The command is short:
bl run loops/price-margin-floor --yes
The control design is the substantive part. Both seed/policy.json and seed/check_margin.py are listed as forbid: paths for the worker. The worker cannot make the gate pass by changing the target margin or weakening the checker. It must repair the catalog to conform to the policy.
The retail recipe protects its policy and checker. The invoice and inventory examples need additional production controls over their reference data before use with live systems.
For a production pricing system, the reference loop is not a deployment guide by itself. A team would still have to supply a real runner, integrate the authoritative catalog source, decide who may approve a repricing, and verify the surrounding controls. The reusable pattern is narrower: keep business policy and evaluation logic outside the worker's authority, then preserve the resulting receipt.
The same structure applies outside retail. A software task can protect its test or lint contract. A compliance workflow can protect a required-clause checker. A research task can protect the source checker that determines whether a cited authority resolves. The property changes. The control relationship should not.
For a broader tour of agent permissions and stopping rules, see the bounded agent harness guide. The project here makes one part of that broader design concrete: the external gate and its receipt.
A gate is only useful when it can fail for the right reason
Putting a checker in a separate file is not enough. The checker itself can be vacuous. It may pass because the artifact it should inspect is absent, malformed, or empty. Or it may apply a rule to a value supplied by the same subject it is supposed to evaluate.
The paper measures gates with destroying mutants. Start from an artifact that a gate accepts, deliberately destroy the property the gate claims to enforce, and observe whether the gate rejects it. This is a focused test of the verifier, rather than a test of whether a worker can produce a pleasing answer.
The paper reports three numbers that need to travel together:
- Across a 69-loop catalogue, the instrument found 47 vacuous gates in shipped, reviewed code.
- After repairs, the evaluated held-out corpus produced 0 false accepts across 209 destroying mutants, reported with a Wilson 95% upper bound of 1.8%.
- When the gates were frozen and a fresh operator family was applied, the paper recovered a 23.3% false-accept rate, reported as 14 false accepts out of 60 in that fresh family.
The third result prevents the first two from becoming a marketing certificate. The reported zero was saturation against the original held-out corpus, not proof that the repaired gates would reject every new failure family. The paper says this directly: a rate belongs to a specific gate, while the apparatus for finding false accepts is the contribution.
The reported rates belong to the evaluated gate families. The fresh post-freeze family exposed false accepts after the earlier corpus reported none.
That is a better operational posture than “our agent passed the tests.” Ask these questions instead:
- What property does the gate actually check?
- Can the worker edit the checker, policy, fixture, or test data that determines the verdict?
- Which destroying changes does the gate reject?
- What credible new failure family has not yet been used to challenge it?
The questions are demanding by design. A gate should not be trusted merely because it is automated or because it appeared to work on familiar examples.
When a loop should become a graph
One bounded loop is often enough: fix a failing test, ensure a record matches a schema, check that a required field exists, or validate a margin floor. The finish condition is local and the result can be explained in one sentence.
A graph earns its complexity when the work has dependencies with different control requirements. The current repository has seven reference graphs, using 28 distinct shipped loop packages. They use a common control skeleton: parallel checks, a join, an approval, one irreversible effect, and a declared remediation branch.
The solo-builder-ship reference graph makes the shape concrete. It runs test-presence, type, and commit-convention checks in parallel. A join waits for their required outcomes. A maintainer approval comes before the publish-instruction node. That final node declares the single external_write effect. A failure-conditioned edge can route a failed test-presence check into a red-green repair loop.
Each loop package is pinned by content digest. The graph therefore names exact package bytes, rather than relying on a mutable package name that could drift after the graph was reviewed. The repository's reference-graph test detects a changed package digest and points maintainers to the regeneration step.
This diagram traces the checked-in solo-builder-ship manifest; it is a contract view, not a new execution receipt.
Archived local-run replay: three successful checks, join, approval pause, and local effect receipt. The raw event log is not distributed in the repository.
Actual monitor screenshot from the repository. It shows the graph on screen; the animated replay above is a separate explanatory asset.
There is one limit worth understanding before building repair paths. A downstream verification node may discover that an upstream node needs another attempt. If each node independently owns an unbounded repair counter, nodes can repeatedly reopen one another. The graph model uses a global repair_budget for the run. Under the stated bounded-repair semantics, total node executions are bounded by:
(1 + repair_budget) × Σ max_attempts(node)
This is a constrained claim. It applies to the declared graph semantics, global repair budget, and per-node attempt limits. It does not say that every graph a team writes will terminate by magic. The point of the contract is to expose the conditions under which “try again” remains finite.
A graph manifest is a control contract
The graph is more than a visualization of agent steps. It states the allowed execution relationships:
- which checks can run in parallel;
- which predecessors must complete before a join can proceed;
- where a human decision is required;
- which branch is allowed to respond to a declared failure state;
- what effect a node may have; and
- what isolation, data class, connection capability, and budget each node declares.
The current graph-capabilities documentation distinguishes shipped behavior from deployer-provided seams. It describes strict manifest validation, connection admission, conditional edges, approval pauses, hash-chained replay, and host-dependent execution modes. That boundary should be retained in any serious adoption review.
For example, a graph that plans successfully has not necessarily run on a host that can meet every declared isolation requirement. A polished monitor view is not, by itself, evidence of a safely executed production workflow. The execution environment, admission checks, human approvals, and actual receipt all matter.
The repository's graph guide is equally specific about the reference examples: the keyless graphs use stub workers and mechanical gates; end-to-end execution requires a host that can enforce each node's declared isolation; and explicit approval tests cover resume-and-publish control paths. These are healthy boundaries to state upfront. They stop an attractive diagram from becoming an unsupported operational claim.
What a real receipt can and cannot tell you
The repository includes a documented Codex receipt dated 13 July 2026. In that run, codex-cli 0.144.3 worked on a citation-existence task. The source brief began with two invalid citations. The documented result says Codex corrected one citation and removed a fabricated authority in one turn without editing the protected reporter or checker. The independent command gate exited 0.
The persisted ledger records a DONE decision, one lap, 228,149 reported input/output tokens, and 68.93 seconds of wall-clock time. Those values are useful because they are tied to a named task, run ID, gate command, and receipt. They are not a benchmark for all Codex work, a promise about price, or evidence that a different model will behave the same way.
This is where receipts change a review conversation. Instead of asking someone to trust a summary, a reviewer can inspect:
- the task and gate that were actually run;
- the terminal decision and number of laps;
- the recorded timing and reported usage;
- the protected paths and relevant artifacts; and
- the scope of any redaction applied to publicly shareable evidence.
An archive can still need qualification. The launch package includes an explanatory replay based on an archived local solo-builder-ship graph execution: parallel checks, join, approval, and a local effect receipt. The raw controller log is not distributed in the repository, so a reader cannot independently verify that replay from public artifacts. It also lacks a clear software-version attestation that would let a viewer claim today's source produced the identical run.
Cite the research
The arXiv paper gives the formal and empirical foundation for Bounded Loops. The repository supplies the runnable implementation, documentation, examples, and tests.
For citation, use the paper's current arXiv record:
@article{260927871,
title={{Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses}},
author={Varun Pratap Bhardwaj and Garima Singh and Arun Pratap Bhardwaj},
year={{2026}},
eprint={{2609.27871}},
archivePrefix={{arXiv}}
}
How to adopt the pattern without turning it into process theatre
Start with one task whose completion property you can explain in a sentence. A test must pass. A document must contain a checked set of clauses. A catalog must meet a fixed margin floor. A response must contain citations that resolve against a protected reporter.
Then build the smallest defensible control loop:
- Put the acceptance rule in an inspectable gate.
- Keep the gate and its policy inputs outside the worker's write authority.
- Declare attempt, time, spend, progress, and approval limits before the run.
- Keep the terminal rule in the controller: a passing gate can yield
DONE; a worker's completion statement cannot. - Store a receipt and review it after the run.
- Attack the gate with destroying changes, including fresh failure families after it appears reliable.
Only add a graph when the workflow truly has parallel checks, joins, explicit approvals, bounded repair, or a material effect that deserves its own release point. A graph is a useful control contract when the dependencies are real. It is theatre when it turns a simple local check into a diagram with no independently meaningful gate.
For teams using coding agents, that yields a practical decision rule: let the model work where it is useful, but move the authority to declare completion into a check the model cannot rewrite or narrate into existence. Then make the cost of further work and the evidence of the finished work inspectable.
Models will improve. That will not remove the need to specify what success means, what a worker may change, how much authority the run has, and who can verify the outcome.
Production links and source notes
- Repository and current documentation: qualixar/bounded-loops
- Research paper: arXiv:2609.27871
- Quickstart: README
- Independent-gate control path: architecture documentation
- Gate-mutation results and research scope: paper abstract
- Retail policy example: price-margin-floor
- Reference graphs: graphs directory
- Graph execution boundaries and repair semantics: graph capabilities
- Documented real run: Codex receipt example









Top comments (0)