I use Claude Code most of the day. Most turns are small. Read a file. Run a test. Rename a thing. Some turns are hard. A design question. A bug that makes no sense.
Running all of them on the biggest model costs money and time. Switching by hand with /model works. For about a day. Then you forget.
So I wrote Router. It's a Claude Code plugin. For each turn of the main conversation it picks the model and the effort level.
IMAGE: router-band.png (the line above the prompt)
What it does
A small classifier gives each new turn a label: micro, low, medium or high. Each label maps to a model and an effort:
| Tier | Model | Effort |
|---|---|---|
| micro | Haiku 5.5 | medium |
| low | Haiku 5.5 | high |
| medium | Opus 5.5 | medium |
| high | Opus 5.5 | xhigh |
Tool calls in the same turn keep the label. Subagents keep their own model.
No proxy
Many "routers" for Claude Code sit between the CLI and the API. This one doesn't. Claude Code has a hook API for plugins. They call them Mods. Router only changes the model and effort of the next step.
Claude Code still builds the request, holds your keys, streams the answer, runs the tools and counts the cost. No local server. No port. No base URL to set and later forget to unset.
The hard part: when not to switch
Switching models is not free. The new model starts with a cold prompt cache. The next request pays to write the whole context into it. On a 300K token session that can cost more than the cheaper model saves.
So before it switches, the router:
- prices the cache write on the new model against the best case for the current one,
- wants a few agreeing turns before it moves to a cheaper model,
- checks with a real measurement that the context fits a smaller window, and keeps 20% free,
- stops for the rest of the turn if Claude Code falls back to another model.
A router that flips models every turn is just a more expensive way to be confused.
Why these defaults
They come from Anthropic's own effort charts. On OSWorld 2.1, Haiku 5.5 at high effort scored above Sonnet 5.5 at low, for about a quarter of the cost. On FrontierCode and CursorBench, Opus 5.5 at medium matched or beat Sonnet 5.5 at xhigh, for 25 to 50% less per task.
Benchmarks are benchmarks. These are defaults, not a promise. You can change every tier.
You can see what it does
A line above the prompt shows the tier, the model, the reason, the context and the cache. /router opens a pane. There you pin a tier for the next turn, change the setup with Save and Undo, tune the policy and pick the classifier.
IMAGE: router-pane-now.png (the Now tab)
What it sends, and when it gives up
The classifier is one you pick: Jev on typesafe.ai, Clef / Clef Flash on Cloudflare Workers AI, OpenAI's gpt-6-luna, or a local Ollama model. The hosted ones get the current prompt and up to 6 recent messages, 1,200 characters each. Ollama keeps them on your machine. No tool output. No system prompt.
No key, a timeout or a network error: the turn keeps its model. Nothing blocks you.
Does it save money?
Probably. How much? I don't know yet. The pane shows Claude's own usage and list prices. I'd rather measure than make up a nice number for this post.
Try it
You need Claude Code 2.1.289 or newer and a key for one hosted classifier, or a local Ollama model.
claude plugin marketplace add alexei-led/claude-router
claude plugin install router@alexei-led-claude-router
Then run /plugin configure router and save the key.
The code is on GitHub, MIT: https://github.com/alexei-led/claude-router
If it picks wrong on a turn, open an issue with the kind of prompt. That helps me most.
Top comments (0)