DEV Community

Solon Framework
Solon Framework

Posted on

An Agent That Finishes: Inside Solon AI's Loop Engine

Most AI coding agents fail in the same way. Not with a wrong answer — with an incomplete one. They start confidently, get 70% of the way there, announce "done!", and hand you a task you now have to finish yourself.

That failure isn't a prompting problem. It's a control-flow problem: nothing in the system knows how to tell "finished" from "tired".

Solon AI ships a module whose entire job is that distinction — solon-ai-loop.

Repo: https://github.com/opensolon/solon-ai (module: solon-ai-loop)

Not to be confused with the /loop command in SolonCode. That's a feature of a product. This is a Java library you embed in your own application.


1. What it actually is

Strip away the naming and solon-ai-loop is four things bolted together:

  1. A state machine — where the work is, and which transitions are legal.
  2. A strategy — what one iteration means (implement a story? run a pipeline phase? run the test gate?).
  3. A validator — whether this iteration passed, and if not, why.
  4. Durable state — so the loop can survive a restart, a crash, or a human going home.

The point of the module is that none of these is left implicit. A loop ends because a criterion was met, or because a patience budget ran out — never because the model stopped talking.


2. The state machine

Eight states, with a whitelist of legal transitions:

State Active Terminal Pausable
IDLE ✗ ✗ ✗
PLANNING ✓ ✗ ✓
EXECUTING ✓ ✗ ✓
VERIFYING ✓ ✗ ✓
FIXING ✓ ✗ ✓
PAUSED ✗ ✗ — (resumable)
COMPLETED ✗ ✓ ✗
FAILED ✗ ✓ ✗
IDLE → PLANNING → EXECUTING → VERIFYING → COMPLETED
                       ↑            │
                       └── FIXING ◄─┘      (the fix loop)
Enter fullscreen mode Exit fullscreen mode

The transition table is enforced in code, not in comments:

From Allowed to
IDLE PLANNING
PLANNING EXECUTING, PAUSED, FAILED
EXECUTING VERIFYING, PAUSED, FAILED
VERIFYING COMPLETED, FIXING, PAUSED, FAILED
FIXING EXECUTING, PAUSED, FAILED
PAUSED PLANNING, EXECUTING, VERIFYING, FIXING, FAILED
COMPLETED / FAILED — (terminal)

Two things worth noticing. First, FIXING can only go back to EXECUTING — a fix is work, not a shortcut to done. Second, VERIFYING is the only state that can reach COMPLETED. There is no code path where an executor declares victory about its own output.


3. Three strategies, three shapes of "keep going"

Ralph — PRD-driven story loop

Reads a PRD, takes the next unfinished user story by priority, implements it, verifies it, records progress, repeats.

RalphLoopStrategy strategy = RalphLoopStrategy.builder()
    .verificationRequired(true)
    .criticMode("architect")        // architect / critic / codex / none
    .maxIterations(50)
    .storyImplementor((task, ctx) -> { /* your agent goes here */ })
    .storyValidator((task, result, ctx) -> /* Boolean */ true)
    .build();
Enter fullscreen mode Exit fullscreen mode

The two hooks are plain functional interfaces — StoryImplementor is a BiFunction<String, LoopContext, Object>, StoryValidator a TriFunction<String, Object, LoopContext, Boolean>. So the loop engine doesn't care how a story gets implemented. Wire it to a ReActAgent, a shell command, a human, whatever. The loop only owns the rhythm.

Team Pipeline — phase-ordered collaboration

Runs a fixed sequence of phases with guards between them:

TeamPipelineStrategy.builder()
    .phases(Arrays.asList(Phase.PLAN, Phase.PRD, Phase.EXEC, Phase.VERIFY, Phase.FIX))
    .maxFixAttempts(3)
    .build();
Enter fullscreen mode Exit fullscreen mode

Phases: PLAN, PRD, EXEC, VERIFY, FIX, plus COMPLETED / FAILED / CANCELLED. The VERIFY phase won't proceed unless tasksCompleted >= tasksTotal, and the fix loop gives up after maxFixAttempts instead of oscillating forever.

UltraQA — the quality gate loop

Run build / test / lint / typecheck. If it fails, fix and run again. Repeat until green, or until you name why you stopped.

UltraQAStrategy.builder()
    .goalType(UltraQAStrategy.UltraQAGoalType.TESTS)   // TESTS / BUILD / LINT / TYPECHECK / CUSTOM
    .maxTestAttempts(10)
    .build();
Enter fullscreen mode Exit fullscreen mode

4. The part I actually like: named exit reasons

GOAL_MET  · MAX_CYCLES · SAME_FAILURE · ENV_ERROR · CANCELLED
Enter fullscreen mode Exit fullscreen mode

SAME_FAILURE is the interesting one. Every failed gate run is normalised — timestamps, line numbers and other noise stripped — then compared. If the same failure shows up 3 times in a row (that's SAME_FAILURE_THRESHOLD, a public constant), the loop stops.

That's a very old idea from build systems, and it's exactly right for agents: a loop that keeps producing the identical error is not making progress, no matter how busy it looks. Killing it early and loudly is the feature.


5. State that outlives the JVM

Long tasks and short processes are a bad match. So the engine can persist to disk:

.solon-ai-loop/
├── state/
│   ├── ralph/{sessionId}.json         # Ralph state (+ strategy mutex)
│   ├── team/{sessionId}.json          # Team Pipeline state
│   ├── ultraqa/{sessionId}.json       # UltraQA state
│   └── sessions/{sessionId}.json      # session index, for querying
├── prd/{sessionId}.json               # the PRD document
└── progress/{sessionId}.txt           # progress memory
Enter fullscreen mode Exit fullscreen mode

Each file is wrapped in two layers — _meta (written_at, mode, sessionId) and data (the full state). Writes go through an atomic-write helper, and the base directory is created with 0700 permissions. The sessions/ index is what makes "what was running when I killed it?" an answerable question.

Because the three strategies share a state directory, they also share a mutex: MutualExclusionGuard refuses to start Ralph while UltraQA holds the lock (canStartRalph / canStartUltraQA / canStartTeam, with stale-lock cleanup). Two loops fighting over one workspace is a nasty failure mode, and it's designed out rather than documented away.


6. Validation is an interface, not an opinion

public interface Validator {
    ValidationResult validate(Object result, ValidationCriteria criteria);
    ValidationResult validateQualityGate(QualityGate gate, Object result);
    ValidationResult validateIteration(Object iterationResult, ValidationContext context);
}
Enter fullscreen mode Exit fullscreen mode

Results are one of three things — passed(message), failed(message, details), or needsFix(message, errors). That third one is what feeds the FIXING state.

Preset gates cover the boring 80%:

QualityGate.build();   // compilation, dependencies
QualityGate.test();    // unit-tests, integration-tests
QualityGate.lint();    // style, complexity, duplication
Enter fullscreen mode Exit fullscreen mode

And the module ships two built-in verifiers, ArchitectVerifier and CriticVerifier, the latter with three modes — architect (architecture-level changes), critic (general review) and codex (test coverage, null handling).

Internally, each verification carries its own little state machine: PENDING → IMPLEMENTED → AWAITING_REVIEW → ARCHITECT_APPROVED → CRITIC_APPROVED, with FAILED after the attempt budget (3 by default) and SKIPPED as the escape hatch. Verification that isn't tracked is just an opinion; this one leaves a trail.


7. Autopilot: chaining strategies into one pipeline

The AutopilotExecutor composes the whole thing into five stages, each bound to a default strategy:

Stage Default strategy
EXPANSION — requirement analysis Team Pipeline
PLANNING Team Pipeline
EXECUTION Ralph
QA UltraQA
VALIDATION Team Pipeline
COMPLETED / FAILED —
PipelineConfig config = PipelineConfig.builder()
    .expansionEnabled(true).planningEnabled(true)
    .executionEnabled(true).qaEnabled(true).validationEnabled(true)
    .strategyForStage(PipelineStage.EXECUTION, RalphLoopStrategy.builder().maxIterations(20).build())
    .strategyForStage(PipelineStage.QA,        UltraQAStrategy.builder().maxTestAttempts(5).build())
    .build();

AutopilotExecutor autopilot = new AutopilotExecutor(engine, config);
var future = autopilot.startPipeline(
    AutopilotExecutor.PipelineRequest.create(sessionId, "Build feature X"));

System.out.println(autopilot.formatPipelineHUD(sessionId));
Enter fullscreen mode Exit fullscreen mode

Any stage can be replaced via the StageAdapter SPI, and the pipeline can be driven by hand — skipStage, advanceStage, cancelPipeline. That matters more than it sounds: an autonomous pipeline you can't interrupt is a liability, and one you can't inspect is a black box.


8. Wiring it into a Solon app

The module has first-class integration with three neighbours — solon-ai-agent (agents drive iterations), solon-flow (a flow context drives phases), and solon-ai-harness (tool management drives the QA loop):

LoopAutoConfiguration.IntegratedComponents c = LoopAutoConfiguration.createDefault();

LoopEngine engine = c.loopEngine;
// c.agentIntegration, c.flowIntegration, c.harnessIntegration

SimpleAgent agent = SimpleAgent.of().chatModel(chatModel).build();
LoopSession session = c.agentIntegration
    .startAgentDrivenRalphLoop("Implement user management", agent);
Enter fullscreen mode Exit fullscreen mode

9. Getting started

<dependency>
    <groupId>org.noear</groupId>
    <artifactId>solon-ai-loop</artifactId>
    <version>4.1.0</version>
</dependency>
Enter fullscreen mode Exit fullscreen mode

Minimal run:

// one-liner, in-memory
LoopEngine engine = LoopAutoConfiguration.createDefaultEngine();

// or durable: state under the project dir, monitoring on
LoopEngine engine = new LoopAutoConfiguration()
    .useDiskState("/path/to/project")
    .enableMonitoring(true)
    .build();

LoopConfig config = LoopConfig.builder()
    .taskDescription("Implement user login feature")
    .strategy(RalphLoopStrategy.builder().verificationRequired(false).maxIterations(5).build())
    .maxIterations(5)
    .build();

LoopSession session = engine.start(config);
session.onStateChange(s -> System.out.println("state: " + s));
session.waitForCompletion(Duration.ofSeconds(10));

LoopResult result = session.getResult();
System.out.println(result.isSuccess() + " / iterations: " + result.getTotalIterations());
Enter fullscreen mode Exit fullscreen mode

10. When not to use this

Being fair about scope:

  • Don't reach for it for a single tool call. If a ChatModel + one tool answers the question, a loop engine is overhead.
  • It does not implement anything for you. You supply the implementor, the validator, or the stage adapter. It provides the rhythm and the memory, not the intelligence.
  • Disk persistence is local. The state manager writes to the project's filesystem — it's not a distributed job queue. Multi-node orchestration is a different problem.

What it removes is the boring, failure-prone part: deciding what "done" means, remembering where you were, and being honest about when to stop. Every agent framework eventually grows this. Solon AI's version is small, explicit, and — in the spirit of the framework — named after exactly one thing it does.


Verified against the solon-ai-loop module source (64 Java files) and Maven Central metadata on 2026-10-06. Code samples are taken from the module's own APIs; the version shown is the current 4.1.x line.

Top comments (0)