DEV Community

Aamer Mihaysi
Aamer Mihaysi

Posted on

If every layer prefix is a valid model, why do we still pick a size at deploy time?

I keep four checkpoints of the same family on disk: a 1.5B, an 8B, a 32B, and a 70B. Four training runs I didn't do, four eval suites I have to trust on faith, four quantized copies, four latency profiles, four rows in the cost table. And at request time I still can't answer the only question that actually matters: how much model does this request need?

Most of my traffic doesn't need the 70B. I know that because I've A/B'd it. The extraction and classification jobs are indistinguishable between the 8B and the 70B, and the multi-step planning traces fall apart below 32B. So I route by hand, using a rules table I wrote at 1am, and I re-verify the whole thing every time a model version bumps. It's the least glamorous part of my job and it eats more time than the agent code does.

Telescopic Language Models argues that this arrangement is a packaging artifact. The claim, roughly: train once with an objective that keeps every layer prefix valid, and you get a usable model at layer 12, at layer 24, at layer 36, at layer 48. Not a truncated model. Not a base model with an exit head bolted on after the fact. A model that was trained to be good at that depth.

If that holds, the small/medium/large release split stops being a capability boundary and becomes a distribution decision. One artifact, N depths, and the serving layer picks per request.

So why do we still commit to a size at deploy time?

The reasons we pin a size are real, not lazy

It's tempting to say we do it out of habit. It's mostly not habit.

Capacity planning assumes a fixed cost per replica. If depth varies per request, your p99 stops being a number and becomes a distribution, and autoscalers are bad at distributions. You can't tell a Kubernetes HPA "scale when the average exit depth crosses 30" without writing a custom metric and hoping it doesn't oscillate.

Batching is worse. Continuous batching runs a batch to the step count of its slowest member. If half the batch exits at layer 16 and half runs to 48, the shallow half didn't save you anything — you burned 32 extra layers of compute on tokens that were already done. To get the savings you need depth-aware scheduling, which means sorting the queue by predicted depth, which adds queueing latency to the requests you were trying to make fast. That's a genuine trade, not a detail.

Then there's the boring stuff. Evals multiply: N prefixes means N regression surfaces, and "which model served this request" stops being a constant you can put in a bug report. Pricing pages want a named thing. Procurement wants a named thing. "Depth 22 of 48" is a hard sentence to sell to someone signing a three-year contract.

And the failure mode is silent. A timeout is loud. A request served at depth 12 that quietly gave a worse answer is invisible unless you're measuring quality per depth per request class, which almost nobody is.

What I'd actually try first

Not a learned router. Learned routers are how you get a system that's confidently wrong in a way you can't reproduce.

I'd start with a static policy keyed on request class, because I already know the shape of my traffic. JSON extraction, classification, short tool-call formatting: shallow. Code generation, multi-step planning, long-context synthesis: full depth. Then I'd measure the quality delta on my own traces — not on MMLU, on the messy half-English tool-call logs that actually hit my endpoints. If the delta on the shallow class is inside noise, the policy ships. If it isn't, I've learned something about where the prefix stops being valid, which is useful either way.

The thing I'd try before any of that, though, is self-speculative decoding. If the shallow prefix is a valid model, it's a free draft model for its own deeper self — same weights, same tokenizer, same training run, nothing to keep in sync. You draft at depth 16, verify at depth 48, and the output is exactly what the full model would have produced. That's a latency win with a correctness guarantee attached, which is rare enough that I'd take it even if the per-request routing idea never pans out. On my box the 8B at Q4 gives me roughly 70 tok/s and the 70B on two rented A100s sits around 20, so the gap I'd be trying to close is real and expensive.

Only after that would I consider a difficulty predictor, and I'd want it conservative — default to full depth, short-circuit only on high confidence, and log every short-circuit decision so I can audit the ones that went wrong.

The part I'm not sure about

I haven't run this. The claim I'd want to verify on my own data is "valid at every prefix," because valid on a benchmark suite and valid on my traces are different claims. My guess is that shallow prefixes hold up fine on format-following and fall apart on anything requiring a chain of inference — which is exactly the shape that makes a static policy workable and a learned router dangerous. If the shallow prefixes are good at structure and bad at reasoning, you don't need a predictor at all. You need to know which of your endpoints are structure problems.

There's also a version of this that's just distillation with extra steps, and I'd want to see the comparison against a properly tuned distilled student before I believe the telescopic framing buys anything. Distillation gives you a fixed small model. This gives you a continuum. Whether the continuum is worth the scheduling complexity is an empirical question, and I don't think anyone has answered it yet.

So, back to the question. Why do we still commit to a size at deploy time? Because a decade of serving infrastructure, autoscaling, pricing and eval tooling was built on the assumption that a model is one artifact with one cost. Telescopic training breaks the assumption. It doesn't break the tooling. Someone has to rewrite the scheduler, the autoscaler and the eval harness before any of this reaches prod.

My bet is the first real deployment isn't a chatbot. It's a batch pipeline where latency doesn't matter and someone can just turn the depth down for the easy half of the queue. Unglamorous, and exactly the kind of place where this actually pays for itself.

Top comments (0)