Most production model routers work at the prompt boundary. You send a request to Cursor or RouteLLM, a classifier guesses whether the task is hard, and one model generates the entire response from the first token to the stop sequence.
Algorithmic papers over the last year have pushed a finer granularity. Even inside a hard reasoning prompt, most generated tokens are boilerplate. Variable names, syntax glue, standard imports, and connective phrases do not need a 32B or 70B parameter model. Under algorithms like R2R, a 1.5B model can decode 95% of the tokens and hand off only the 5% of high-entropy branching tokens to a 32B model, matching the 32B model's accuracy on AIME math problems. By contrast, a query-level router has to send roughly 40% of whole prompts to the 32B model to hit that same quality bar.
With token savings that lopsided, you would expect everyone to run token-level routing in production. Almost nobody does, because standard inference engines like vLLM and SGLang assume one request belongs to one model. When you wire two separate servers together and bounce a request back and forth every few tokens, throughput craters.
A paper released yesterday by researchers at Tsinghua University and Carnegie Mellon University, TokenRouter: Efficient Serving System for Token-Level LLM Routing, profiles where the time goes and builds a serving runtime on top of SGLang 0.5.1 that gets 2.01x to 64.15x higher decoding throughput than existing implementations.
Where the milliseconds go when you bounce between servers
Suppose you stand up two standard SGLang servers, one hosting Qwen3-0.6B and one hosting Qwen3-32B, with an external Python dispatcher passing tokens between them.
Every time the 0.6B model hits an uncertain token, it pauses that request, sends the context over to the 32B server for one token, and waits for the result to come back. When the request returns to the 0.6B server to continue decoding, the standard server has no concept of a paused request returning from a peer. It treats the returning request as a fresh extend call from a client.
The TokenRouter authors profiled the final 128 decoding steps of a 2,048-token generation on the small-model server under this standard multi-server setup. The per-step breakdown is brutal.
Small-model inference took 1.56 ms, just 4.20% of the step latency. Updating the RadixCache took 12.12 ms (32.62%). Prefix matching took 7.78 ms (20.94%). Locking cache nodes took 6.49 ms (17.47%), and releasing those locks took another 5.13 ms (13.81%).
More than 95% of the runtime on the small model went to radix tree lock contention and prefix matching for a request that was already sitting in GPU memory five milliseconds earlier. When the maximum output length grows from 2,048 tokens to 8,192 tokens, standard serving loses 58.1% to 85.2% of its throughput just re-walking prefixes.
Three assumptions that break when models share a single response
Beyond cache thrashing, token-level routing breaks three assumptions baked into single-LLM serving engines.
First, step synchronization falls apart. In a single-model batch, every active request finishes a forward pass at the same millisecond. When you mix a 0.6B model (6.0 ms per step) and a 32B model (27.9 ms per step), synchronous step barriers force the small model to sit idle for 22 ms every step waiting on the big model.
Second, continuous batching suffers from admission delays. Speculative decoding uses two models too, but the draft-then-verify pattern is fixed and periodic. Token-level routing is data-dependent. A request on the 0.6B model might run for 30 tokens before hitting high entropy, or it might switch twice in four tokens. When a routed request arrives at the 32B subserver mid-batch, it sits in a queue until the current 27.9 ms forward pass finishes. That creates fragmented micro-batches and GPU bubbles.
Third, CUDA graph capture misses the hot path. Standard engines capture CUDA graphs for decode batches, where every request appends a single token. Under token-level routing, every returning handoff arrives with a suffix of new tokens generated by the peer model, which triggers extend mode instead of decode mode. Without CUDA graphs covering variable-length extend batches, kernel launch overhead dominates the handoff path.
Request-centric code, model-centric execution
TokenRouter fixes this by separating the router logic from the batch scheduler.
From the developer's side, you write sequential Python tracing a single request through three methods (route, send, and receive). Here is the small-model router for R-Stitch, which delegates tokens whenever normalized Shannon entropy crosses a threshold:
class RStitchSLM(BaseRouter):
def route(self, result):
probs = result.logits_output.next_token_logits.softmax(dim=-1)
entropy = -(probs * probs.log()).sum(-1) / math.log(probs.shape[-1])
destination = torch.zeros_like(entropy, dtype=torch.long)
destination[entropy > self.tau] = self.peer_idx
return destination.cpu()
def send(self, req) -> PeerReq:
token_ids = req.origin_input_ids + req.output_ids[:-1]
peer_anchor = req.peer_locs[self.peer_idx] or 0
return PeerReq(
rid=req.rid,
token_ids=token_ids[peer_anchor:],
loc=len(token_ids),
finished=req.finished(),
)
That route method runs in-process on the same GPU as the model runner, avoiding a cross-process IPC round-trip on every generated token. Look at peer_anchor in send too. Each request tracks the last token index each peer model saw, so ZeroMQ only ships the unseen token suffix across subservers.
Under the hood, each model runs as an autonomous subserver with three decoupled loops (client-server, local decoding, and inter-model ZeroMQ messaging). When route sends a request to the 32B peer, the 0.6B scheduler keeps the request in memory and flips its state to pending.
While pending, the local decoding loop skips the request, leaving its KV-cache slot and RadixCache pointers pinned in place. When the 32B model sends its token back, the 0.6B subserver appends the token IDs and flips the status back to running. It skips prefix matching, radix lock acquisition, and KV reallocation entirely.
The runtime also captures CUDA graphs over a range of small extend batch sizes and token lengths, and uses a joint KV-cache allocator so two models co-located on shared GPUs via CUDA MPS do not run out of memory when concurrency spikes.
Waiting on purpose with delayed batching
Fixing the handoff overhead gets you partway there. On R2R at concurrency 8 (Qwen3-0.6B + Qwen3-32B), adding extend CUDA graphs and in-process routing raises throughput from 132.78 tokens/s to 230.79 tokens/s. Turning on the asynchronous pending handoff pushes it to 296.86 tokens/s.
The remaining bottleneck is batch fragmentation on the 32B model. If the 32B subserver fires a 27.9 ms forward pass the instant a single routed request arrives, the next three requests that trickle in 4 ms later sit locked out of the batch.
So TokenRouter adds a delayed-batching scheduler. Each subserver buffers incoming routed requests until the queue hits a threshold $B_i$ before launching a forward pass. In their setup, the fast 0.6B model uses $B_0 = 1$ (run immediately), while the slow 32B model holds until $B_1$ requests accumulate.
If $B_1$ is too small, the 32B model runs tiny batches and stalls later arrivals. If $B_1$ is too large, requests get trapped waiting in the 32B queue and starve the 0.6B model of work. To pick $B^*$ without manual grid search, the authors model the multi-model system as a finite-state Discrete-Time Markov Chain parameterized by concurrency $N$, step latencies $L_i$, and the empirical routing probability $p$, solving for the stationary distribution that maximizes committed tokens per millisecond.
At concurrency 8, turning on that analytical delayed-batching threshold pushes R2R throughput from 296.86 to 372.48 tokens/s, a 2.76x gain over the official R2R implementation (134.89 tokens/s). Under a tighter latency budget comparing concurrency 16 against R2R at concurrency 1, TokenRouter hits 18.58x higher throughput while still running 1.13x faster per user. Across all five tested algorithms (R2R, CITER, R-Stitch, Co-LLM, and Mixture-of-Ensembles) and three workloads including 8k-context SWE-Smith agent trajectories, throughput jumps between 2.01x and 64.15x.
Where this fits in a real stack
One honest constraint sits in Section 6 of the paper. The Markov chain solver assumes the number of tokens generated between model switches follows a geometric distribution, meaning a roughly stationary switching probability $p$. In multi-turn coding agents, entropy comes in bursts, such as fifty easy tokens of JSON formatting followed by ten hard tokens of control-flow logic. Even when the geometric assumption drifts, the pending state preservation and extend-mode CUDA graphs carry most of the latency reduction on their own.
If you only call closed frontier APIs over HTTPS, token-level routing is still off the table because network RTT kills per-token handoffs. If you run open-weight models on your own GPUs, where a Qwen3-0.6B draft worker and a Qwen3-32B reasoner sit on the same node or across RoCE, the TokenRouter repo makes fine-grained routing practical without writing custom C++ scheduler hooks.
Top comments (2)
For anyone on the API side of your last paragraph, the routing unit has to be the whole answer, which moves the trigger from token entropy to a check on the finished output. One thing we measured there that might carry over to token-level routing: escalating can make some answers worse. We sent the same 2,400 tool-calling tasks through a cheap model and a frontier model, three runs each. The frontier model fixed 133 answers the cheap one got wrong, and broke 34 the cheap one got right. 10 of those 34 were right in all three cheap runs and wrong in all three frontier runs, which points at how that model handles that item. The totals hide this, because the two directions partly cancel.
Which makes me curious about the R2R comparison. When the 1.5B plus 32B mix matches the 32B model's AIME accuracy, does it match on the same problems, or does it reach the same total with a different set right and wrong? A per-problem table would show whether the 5% of routed tokens are carrying the hard part, or whether the totals just happen to line up.
That 34-break stat on tool calling matches what happens when a larger model overthinks a simple schema.
On R2R and AIME, neither paper published the paired problem-level diffs, and aggregate accuracy almost certainly hides similar churn. Token-level routing has its own version of negative transfer through prefix conditioning. When the 1.5B model runs the first 80 tokens of a proof at low entropy, the 32B model only gets called after the trajectory is already locked in. If the 1.5B model took a subtle conceptual detour that stayed low-entropy until five steps later, the 32B model has to generate tokens inside a prefix it would never have written itself.
On benchmarks like AIME with only 30 problems, two lucky recoveries cancelling two prefix-trapped failures is enough to match the headline number. Without the per-instance diff matrix, there is no guarantee the hybrid rollout followed the 32B model's reasoning graph.