DEV Community

Zara Menon
Zara Menon

Posted on

Claude Code Router v3: What Changed and How I Set It Up Now

I went back to Claude Code Router this week to point a side project at a cheaper model. My notes from the last time I used it were useless. CCR isn't a small proxy you start and forget anymore. The README now calls it "a local model gateway and control plane for coding agents", and the recommended install is a desktop app.

Most of what ranks for "claude code router" still describes the older shape: run a proxy, aim Claude Code's base URL at it, sort traffic into background, thinking and long-context buckets. The core idea holds. The way you set it up and route traffic doesn't. This is what I worked out from the v3 README and docs. The latest release is v3.1.1, published September 16, 2026, and the repo is around 37.6k stars.

(For context, the side project is a landing page plus a small browser extension. I'm testing Begin for that part, since it builds websites and Chrome extensions from the same prompt. Claude Code and CCR are for the backend code I still write by hand.)

What Claude Code Router is in v3

The short version: CCR runs on your machine and gives your coding agents one stable endpoint. Behind that endpoint you manage providers, models, credentials, routing rules and logs.

Three things surprised me:

It's not just for Claude Code. The supported list now includes Claude Code (CLI and app), Claude Design, Codex, Grok CLI, Kimi CLI, Kilo Code, OpenCode, Pi, ZCode and WorkBuddy. The name stuck. The scope didn't.

Config moved into the app. The settings live in config.sqlite in ~/.claude-code-router (or %APPDATA%\claude-code-router on Windows). The docs warn against editing or copying the live SQLite file while CCR runs. Use Settings → Export data, or stop CCR before a file-level backup. If your muscle memory is "open the config and edit JSON", unlearn it.

You launch the agent through CCR. You don't just set a base URL. You create an "Agent Config" profile, then start Claude Code from CCR. Open Claude Code the normal way and it skips CCR entirely, unless the profile's scope is set to "System default".

Here's how the old mental model maps to v3, as far as I can tell from the docs:

What I used to think about Where it lives in v3
Point Claude Code's API base URL at the proxy An Agent Config profile, launched with ccr "<profile name>" or from the desktop card
A default model plus task buckets A profile default model, plus separate Fable / Opus / Sonnet / Haiku tier overrides
Special-case routing Custom rules on request headers or body, or a Node.js script rule
Hoping a provider stays up Per-rule or default fallback: retry or an ordered model-chain
One API key per provider Credential pools with priority, weight and limits
Guessing what happened Request logs showing request model, resolved provider and resolved model

My setup, step by step

I'm on a headless Linux box most of the time, so I used the npm CLI rather than the Electron app. Both read the same config directory, but the commands differ. The desktop app writes launchers named ccr-app. The npm package gives you ccr. Don't mix them up when copying commands from a profile card.

node --version                     # the CLI needs Node.js 22 or newer
npm install -g @musistudio/claude-code-router
ccr ui --no-open                   # background service, no browser on SSH
Enter fullscreen mode Exit fullscreen mode

ccr ui prints a management URL. Two ports matter. The management UI defaults to 127.0.0.1:3458 and the model gateway to 127.0.0.1:3456. If 3458 is taken, CCR moves to the next free port, so trust the printed URL. That URL contains a ccr_web_token query parameter. The docs say to treat it like a password. Don't paste it in a ticket or a screenshot.

Then, in the UI, in the order the docs suggest:

  1. Providers → Add Provider. Pick a preset or "Other / custom API endpoint", paste the key, and let CCR auto-detect the protocol and models. It supports OpenAI Chat, OpenAI Responses, Anthropic Messages, Gemini Generate and Gemini Interactions. Check Connection sends a real request, so it can use a few tokens.
  2. API Keys → Add API key. This is a separate credential: the key Claude Code uses to talk to CCR. You can set an expiry and local request, token or image limits per minute, hour or day. The management token and client keys are independent.
  3. Routing. Set the default model and any rules or fallbacks.
  4. Server. Confirm the gateway is running.
  5. Agent Config → Add profile → Claude Code. Name it, choose the scope, pick a model.

Then launch it:

ccr "Claude Code - Work"
ccr "Claude Code - Work" cli -- --model sonnet   # args after -- go to the agent
ccr stop                                          # stops the background service
Enter fullscreen mode Exit fullscreen mode

Inside Claude Code, /model lists the models CCR exposes. That's the quickest way to confirm you're going through the gateway. The other is the request log.

For a server where something else supervises the process, the docs recommend ccr serve --no-open in the foreground, with a fixed CCR_WEB_AUTH_TOKEN. Don't also run a background ccr start next to it, or you get two processes fighting over one config.

Routing: tiers first, rules second, scripts last

The tier overrides were the biggest change in how I think about routing. Claude Code picks a model per tier, and a CCR profile lets you override each tier separately. A strong provider model on Opus, a cheap one on Haiku, everything else on the default. Leave a tier empty and Claude Code picks for itself. The value has to be a valid Provider/model (the docs use Moonshot/kimi-k3 as an example) or a Fusion model.

Subagents get their own mechanism. If you add a Description to models on the Models page, CCR injects that list into Claude Code's Agent/Task tool descriptions. When Claude Code spawns a subagent, it starts the prompt with a tag like <CCR-SUBAGENT-MODEL>provider/model</CCR-SUBAGENT-MODEL>. CCR strips the tag and routes that request to the tagged model. No descriptions means no injection, so the Description field doubles as the on/off switch. When it works, the request log shows builtin:claude-code-subagent.

Custom rules come after that. They match on request.header or request.body, run top to bottom, and the first enabled match wins. Most of my "it went to the wrong model" moments came down to rule order.

When one condition isn't enough, there's a Node.js script rule. The file is an async function body, not a module. You get a read-only input object and return a decision. This is the one I use to send oversized requests to a long-context model:

// long-context.js: CCR script rule (function body, no imports/exports)
if (input.summary.hasImage) {
  return null; // not mine, let the next rule decide
}

if (input.tokenCount > 120000) {
  return {
    model: "Moonshot/kimi-k3",
    fallback: {
      mode: "model-chain",
      models: ["Provider/backup-model"],
      retryCount: 0
    }
  };
}

return null; // no match, continue down the list
Enter fullscreen mode Exit fullscreen mode

Some details from the docs worth knowing. input.tokenCount is CCR's estimate and is 0 when unavailable. The timeout is configurable from 10 to 30000 ms. CCR re-reads the file on every run, so edits apply without re-saving the rule. Scripts are fail-open: an exception, timeout or invalid result logs a diagnostic and moves on to the next rule. That's the right default for availability. It also means a broken script fails quietly, so check the logs after any edit.

The parts I'm less excited about

Routing doesn't change model quality. A cheaper model behind Claude Code's interface is still the cheaper model. The interface makes it more pleasant, not smarter.

The surface area has also grown a lot. Fusion, ToolHub, browser automation, Chrome login-state import, AgentClaw bots relaying through Slack, Discord, Telegram and more. All optional, but there's much more to secure than a single proxy port. Upstream provider keys sit in the local data directory, and the docs say to protect it and its backups as sensitive data. If you bind the management service to 0.0.0.0, they ask for a fixed strong token, a firewall or private network, and TLS from a reverse proxy.

Finally, the routing is easy to bypass by accident. If Claude Code isn't launched from CCR and the profile isn't "System default", your requests go straight to Anthropic. The troubleshooting page leads with that question for a reason.

FAQ

Does Claude Code Router still work with the npm CLI, or only the desktop app?
Both. npm install -g @musistudio/claude-code-router gives you ccr, which runs the same gateway and a browser-based UI without Electron. You need Node.js 22+.

What port does Claude Code Router use?
The model gateway defaults to http://127.0.0.1:3456 and the management UI to http://127.0.0.1:3458. The UI moves to the next free port if 3458 is busy.

Why aren't my Claude Code requests going through CCR?
Check three things: the CCR service is running, Claude Code was launched from CCR (not opened directly), and the Agent Config profile is enabled with a scope that covers how you're launching it.

How do I fix "model not found" with Claude Code Router?
The model name has to match in three places: the provider's model list, the model in your routing config, and the model in the Agent Config. Compare all three.

Closing thoughts

Claude Code Router grew from a clever hack into real infrastructure. I'd rather have the request logs, credential pools and explicit fallbacks than the old setup. It just takes an afternoon to relearn. The repo README links to the full docs once you need more than the basics.

As for the side project: the backend goes through Claude Code on a routed model. The marketing site and the extension shell go through Begin, which comes with hosting, sign-in and Stripe payments already connected.

Top comments (0)