DEV Community

Machine coding Master
Machine coding Master

Posted on

Stop Restarting Failed Loops: Resumable ReAct Agents with Spring AI State Checkpointing

Stop Restarting Failed Loops: Resumable ReAct Agents with Spring AI State Checkpointing

Your ReAct agent just burned 45,000 tokens navigating an 8-step enterprise workflow, only to blow up on a transient downstream 503 at step 7. Catching that exception and restarting the prompt chain from step zero isn't just terrible engineering—it is lighting your inference budget on fire.

Want to go deeper? javalld.com — machine coding interview problems with working Java code and full execution traces.

Why Most Developers Get This Wrong

  • Treating agents like stateless HTTP calls: Wrapping ChatClient inside a basic @Retryable loop forces the model to regenerate every prior thought and action when a tool flakes out.
  • Side-effect amnesia: Retrying an agent from scratch without an idempotent journal causes non-idempotent tools (like payment authorization or ticket creation) to double-execute.
  • In-memory state reliance: Storing the Thought-Action-Observation trace purely in JVM memory guarantees unrecoverable drops on container redeploys or out-of-memory kills.

The Right Way

Decouple reasoning from tool execution by persisting an append-only state journal at every step of the ReAct cycle.

  • Implement an explicit StateCheckpointRepository that commits the AgentContext and conversation trace immediately after every tool observation.
  • Attach an idempotency key derived from (agentRunId, stepIndex, toolName) to every external side effect.
  • Resume failed workflows deterministically by rehydrating the checkpointed history into ReActAgent rather than re-prompting the initial goal.

Show Me The Code

Configure a durable, checkpoint-aware ReAct workflow in Spring AI:

@Bean
public AgentExecutor orderRecoveryAgent(ChatClient.Builder builder, VectorCheckpointStore store) {
    return ReActAgent.builder(builder.build())
        .tools(inventoryTool, paymentTool, logisticsTool)
        .checkpointStore(store) // Persists state to PostgreSQL after every step
        .onStepFailure((context, failure) -> {
            log.error("Step {} failed for run {}. Checkpointing for resume.", 
                context.getStepIndex(), context.getRunId(), failure);
            return RecoveryStrategy.SUSPEND; // Freezes state, releases thread
        })
        .idempotencyKeyGenerator(ctx -> ctx.getRunId() + ":" + ctx.getStepIndex())
        .build();
}
Enter fullscreen mode Exit fullscreen mode

Key Takeaways

  • Never restart from step zero: Save intermediate Thought/Action/Observation vectors so transient failures only cost you the failed step's tokens.
  • Enforce tool idempotency: Key all side-effecting operations by agent execution ID and step index to make re-entry safe.
  • Treat agents as distributed state machines: If your framework doesn't survive a JVM crash mid-loop, it is not production-ready.

Top comments (0)