I've shipped INT8 KV caches that looked perfect at 4k context and fell apart at 32k. Same weights, same prompts, same eval harness. The only variable was how long the model had been running. That's the failure mode I keep circling back to when I read quantization papers, and it's why STEPQuant caught my attention: it argues that uniform INT8 on the recurrent state of a delta-rule linear attention model is the wrong default, because not every element of that state carries the same error budget. https://arxiv.org/abs/2609.38169v1
I haven't reproduced their numbers. But the argument matches something I've watched happen in production, so let me lay out the comparison as I'd actually make it on my own rig.
Why uniform INT8 is the default
Uniform per-tensor INT8 is the default because it's the only version that's free. One scale, one zero point, one kernel shape, no branching, no gather, no metadata to carry through the decode loop. On a bandwidth-bound decode step, that matters more than people admit — you're not compute-limited, you're moving bytes, and a clean 8-bit load is about as good as it gets.
For weights, this is fine. Weights are static, they get read once per forward pass, and any error is a fixed bias you can measure offline. For a KV cache it's mostly fine too, because each cached token is read a bounded number of times and then it's either evicted or it isn't. Errors don't compound; they just sit there.
The recurrent state in a linear attention model is a different animal, and this is where I think the paper lands a real point.
The state is a running sum, not a buffer
In a delta-rule model, the state is a matrix that gets updated every token roughly like S ← S(I − kkᵀ) + vkᵀ. Two things follow from that.
First, an error injected at step t doesn't stay put. It gets carried forward and multiplied by the transition at every subsequent step. Second — and this is the part I find genuinely sharp — that transition is a projection. The error decays only in the subspace the model keeps querying. An error that lands in a direction the model never queries again is effectively immortal. It sits in the state forever, contributing nothing useful and quietly biasing every read that touches it.
So "how long does this state element survive" is only half the story. The other half is "which key rows does it get hit by." A row that's constantly overwritten by high-norm keys gets scrubbed clean on a short timescale. A row that's written once and then ignored for 50k tokens is a permanent resident, and it deserves more bits than its neighbor.
That's the whole thesis, and it's a deployment thesis, not a training one.
Four options, ranked by how much I'd trust them
Uniform per-tensor INT8. Fastest, simplest, and the thing I'd reach for first. Its failure mode is dynamic range: if a handful of rows dominate the scale, everything else gets an effective 5 or 6 bits. In a state that's mostly cold with a few hot rows, that's exactly the wrong allocation.
Per-row (per-channel) INT8. Better range handling, still uniform in time. Cheap to implement, still one kernel shape. This is the version I'd actually ship today if I needed something this week.
Lifetime-aware two-tier. Recent state in BF16, older state in INT8. This is the same instinct behind keeping the last N tokens of a KV cache in full precision, and it's the easiest of the "smart" options to reason about. Cost is a second buffer and a merge step.
Lifetime plus row impact. Allocate bits per row based on how often it's read and how long it survives. Best quality per byte, worst kernel. This is what the paper is arguing for, and I believe the analysis. I'm less sure I'd pay for it.
The tax nobody puts in the abstract
Mixed precision inside a recurrent state means the kernel can't be a clean vectorized load anymore. You get gather/scatter, you get branch divergence across rows, and on a single consumer GPU that can easily cost more than the bandwidth you saved. I've been burned by this before: a "smarter" quantized path that did less math and ran slower, because the memory access pattern went from streaming to scattered.
The honest comparison isn't quality-per-byte, it's quality-per-byte-per-millisecond. A scheme that saves 30% of state memory but adds 15% to per-token latency is a loss on any interactive workload. It might be a win on a batch-64 offline job where you're capacity-bound instead of latency-bound. Those are different products and they deserve different answers.
What I'd actually measure
Not short-context perplexity. That's the eval that lied to me about my KV cache for months.
I'd run three things. A needle-in-a-haystack at 64k and 128k, because slow state drift only shows up when the model has to reach back far. A "stale probe" — plant a fact at token 200, run to 100k, then query it, which is the specific thing a permanent error in a cold row would break. And instrumentation: log per-row error norms against a BF16 reference state over time. If the paper's claim is right, that log should show a small set of rows accumulating error and never recovering, and those rows should be the ones with low query frequency.
If that log looks flat, the whole argument collapses and uniform INT8 was fine all along. I'd want to see it before I rewrite a kernel.
The general lesson
Uniform anything is a default, not a decision. We already learned this with KV eviction — not all cached tokens matter equally — and with mixed-precision training, where per-tensor loss scaling beat one global scale. Recurrent state quantization is the same shape of problem with a nastier twist, because the state has memory and the KV cache mostly doesn't.
I'll probably keep shipping uniform INT8 for a while, because it works and it's fast and my time is finite. But I'll stop calling it correct. It's just the thing that hasn't bitten me yet.
Top comments (0)