TL;DR
- MonkeysCode Auto mode is now live across Editor 1.3.1, CLI 1.1.1, and Agent Manager 1.1.1.
- The router dynamically shifts models across distinct agent phases (
plan→execute→review). - Context handoff briefs bridge state between frontier planning models and standard execution models without polluting context.
- Pricing is calculated strictly per-turn based on the routed model, avoiding opaque flat premiums.
Developers using AI coding agents face a consistent balancing act between cost and capability. Routing all requests to a frontier model provides high reasoning capacity for architecture and multi-file planning, but burns through budget on routine execution tasks like boilerplate generation or minor syntax edits. Conversely, pinning an agent to a cheaper, smaller model saves money but routinely fails when the agent needs to plan complex state changes across a large repository.
To address this, MonkeysCode has shipped Auto mode across the MonkeysCode Editor, CLI, and Agent Manager. Auto mode replaces static model selection with a dynamic router that shifts models based on the current phase of the agent's workflow.
The Core Problem: Cost vs. Capability
Before Auto mode, developers had to manually toggle between models depending on the task at hand. If you needed Capuchin, our built-in coding agent, to scaffold a new microservice, you would manually select a frontier model. If you just needed it to write unit tests for an existing file, you might switch to a faster, cheaper model.
This manual switching creates friction. Developers often forget to switch back, leading to either failed complex tasks (because a small model was left active) or inflated bills (because a frontier model was used for hours of simple refactoring).
Auto mode removes this friction by making the routing decision programmatically based on the active phase of the agent's lifecycle.
Phase-Based Routing Architecture
Auto mode operates by breaking down a developer's prompt into a strict lifecycle: plan → execute → review. Rather than handling this entirely within a single context window of a single model, the MonkeysCode model-proxy and context-kernel route the work dynamically.
- Plan: The initial request requires deep reasoning. The router assigns this to a frontier model. The frontier model analyzes the codebase, determines the necessary file changes, and constructs a detailed execution plan.
- Execute: Once the plan is approved, the task shifts to the execute phase. The router escalates the execution steps to a faster, more cost-effective model suited for writing the actual code.
- Review: After execution, a review phase validates the changes, ensuring the implementation matches the initial architecture plan.
This role-based routing was engineered over several distinct build steps. In build step 7, we formalized the roles inside Auto (plan → execute → review). In build step 8, we integrated the SDK plan error and route events, and added Auto mode UI elements like the plan card and routing chips to the Agent Manager. Furthermore, as documented in our D14 and D15 specifications, every app opens on Capuchin by default, and any subsequent fix round is treated as a new task, ensuring the router re-evaluates the necessary model for the fix.
Context Handoffs
A major technical hurdle in multi-model routing is maintaining state. Passing the entire planning dialogue into an execution model is inefficient. It wastes context window space, increases latency, and risks confusing smaller models with excessive architectural deliberation.
To solve this, we introduced the handoff brief in build step 5. The model-proxy generates a concise handoff brief that distills the architectural decisions and exact file targets from the plan phase. This brief is passed to the execution model, providing a clean, highly focused state without the overhead of the entire conversational history. The context kernel also provides phase and signal hints directly to the models to ensure they understand their current role in the pipeline.
Transparent Per-Turn Pricing
A common industry practice for "auto" or "premium" modes is to charge a flat, inflated rate regardless of which model actually processes the request. We rejected this approach.
With the release of Editor 1.3.1 and CLI 1.1.1, we explicitly dropped the "Auto frontier" button and removed the /frontier command. Instead, the context-kernel prices Auto turns strictly as the routed model. If the router assigns an execution task to a cheaper model, you are billed the exact rate for that cheaper model.
To make this transparent, we shipped a dedicated Auto usage card and route log in the web dashboard (build step 9), accessible via the GET /v1/usage/auto-summary endpoint. The web dashboard now features a dedicated Auto mode page, a home section highlighting active routes, and detailed billing highlights. You can view your exact per-model breakdown at monkeyscode.com/pricing.
Using Auto Mode
Auto mode is now the default experience across the MonkeysCode ecosystem.
In the CLI
We have updated the MonkeysCode CLI to version 1.1.1. The /model command now prominently displays Auto. We also shipped CLI tabs, mouse and command menus, and an Auto usage card directly in the terminal interface.
# Update to the latest CLI
npm install -g monkeyscode-cli
# Start an agent session in Auto mode
monkeyscode run --mode auto
In addition to the CLI updates, build step 9 introduced Agent Manager SSH support, model-proxy identity verification, and platform CORS updates to support distributed agent deployments.
In the Editor and Agent Manager
In Editor 1.3.1, we introduced the Auto picker, route chips, and plan cards. We quickly followed up with Editor 1.3.2 and Agent Manager 1.1.2, which introduced plan approval under Auto.
This means that even when Auto mode is handling the routing, you retain control. The frontier model generates the plan, and the Editor presents a plan approval UI. Only after you approve the plan does the model-proxy generate the handoff brief and route the task to the execution model. We also added an Auto status strip and a per-tab change tray in 1.3.2 to make the active phase and model completely visible.
For teams managing internal models, Auto mode respects the admin profile. The versioned Auto mode model config and admin API (shipped in build step 4) ensure that Auto only routes to servable models approved by your organization. The platform also keeps the Auto config partial index intact across schema:update operations, ensuring zero downtime for your routing rules.
Takeaway
Auto mode is not magic; it is pragmatic resource allocation. By formally separating the agent lifecycle into plan, execute, and review phases, and utilizing context handoff briefs, MonkeysCode allows developers to leverage the reasoning of frontier models without paying frontier prices for basic syntax edits.
Update to Editor 1.3.2 and CLI 1.1.1 today to start using Auto mode. If you are new to the platform, we recently shipped a permanent 14-day promo trial (replacing the old referral system)—check out the MonkeysCode documentation to get started.
Top comments (0)