DEV Community

Alex Fank
Alex Fank

Posted on

Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term

llmrun.dev estimates whether a given LLM can be run locally and how fast it will decode. The first part of that question is arithmetic. The second required several iterations, and Mixture-of-Experts models were the case where the initial formula was least accurate. This note describes the estimator in its current form, the calibration behind it, and the cases in which it is known to fail.

Decode is a bandwidth problem

During generation, each output token requires reading the weights from memory once. Batch size is 1 on a local machine, so there is almost no arithmetic over which to amortize that read, and speed is set by how fast bytes stream out of VRAM or unified memory. The dense formula used is the standard one:

tok/s = (bandwidth GB/s / model size GB) x efficiency
Enter fullscreen mode Exit fullscreen mode

Efficiency is a per-platform constant covering software overhead and the gap between spec-sheet and achievable bandwidth. In the code it is 0.65 for NVIDIA (CUDA), 0.60 for AMD (ROCm 7.x), 0.50 for Intel (oneAPI) and 0.70 for Apple (Metal). These values were taken from community llama.cpp and Ollama results and are treated as calibration constants, not physical quantities. The basis of the estimate is dated in the code (last reviewed 2026-09-21, against llama.cpp / Ollama with GGUF weights), and the methodology page prints that date.

Why dense math fails for MoE

A Mixture-of-Experts model holds all its experts in memory but routes each token through only a few. VRAM therefore scales with total parameters, while bytes read per token scale with active parameters. gpt-oss-20b has 20.9B total parameters and 3.6B active. Applying the dense formula to its 12.9 GB footprint on an RTX 4090 (1008 GB/s, from the spec sheet) gives about 51 tok/s, whereas llama.cpp's llama-bench measures about 222 on that card. The dense formula is off by more than 4x, in the pessimistic direction.

The first correction is to scale the bytes read by the active fraction: the in-memory model size is multiplied by active/total parameters. This is an approximation, since attention and shared weights are not routed and the active count already includes them, but it allows a single ratio taken from the model's config to be used instead of modelling tensors separately. A model takes this path only if active parameters are below 90% of total. Anything denser is treated as dense, in which case the function returns exactly the dense result.

The routing term

Scaling bytes alone overshoots. When only 3.6B parameters are read per token, the bandwidth term becomes very small, and a fixed per-layer cost that the formula ignored begins to dominate: evaluating the router and dispatching to expert kernels, once per layer, per token. This cost is modelled as an additive time per layer:

seconds/token = active_GB / (bandwidth x efficiency)
              + layers x per_layer_overhead_ms / 1000
Enter fullscreen mode Exit fullscreen mode

The per-layer overhead constants are 0.055 ms for NVIDIA and 0.2 ms for Apple, AMD and Intel, with 0.1 ms when the brand is unknown. The NVIDIA value is much lower because CUDA graphs amortize launch cost better than the other llama.cpp backends do. If the layer count is missing, a fallback of 48 is used, the median across MoE models in the catalogue at the time of calibration. The constants were fitted against llama-bench tg128 (decode) results for gpt-oss-20b and gpt-oss-120b on 8 hardware combinations. They are kept in the test suite as fixtures, so a change to any constant that breaks one of them fails the build.

Two worked examples follow, computed from the constants above. The bandwidth values are the GPU/SoC spec values used in the fixtures. The measured figures are the llama-bench results in the fixtures.

RTX 4090, gpt-oss-20b M2 Ultra, gpt-oss-120b
Model in memory 12.9 GB 65 GB
Active / total params 3.6B / 20.9B 5.1B / 116.8B
Bytes read per token 2.22 GB 2.84 GB
Bandwidth (spec) x efficiency 1008 x 0.65 800 x 0.70
Bandwidth term 3.39 ms 5.07 ms
Layers x overhead 24 x 0.055 = 1.32 ms 36 x 0.2 = 7.20 ms
Estimate 212 tok/s 81.5 tok/s
Measured (llama-bench) 222 80

The 4090 case is dominated by bandwidth. In the M2 Ultra case, routing overhead accounts for more than half of the per-token time.

Accuracy

The test suite asserts that every fixture lies within 18% of its measured value. The public page states an expected error of roughly plus or minus 20% on decode tok/s, and quoting a tighter figure is not justified. Eight fixtures on two model families constitute a small sample. The constants fit gpt-oss well. They have not been validated on every MoE shape (many fine-grained experts, large shared experts), and misses of more than 20% are expected for some of them.

Where it breaks

CPU and RAM offload is the largest gap. Once part of the model resides in system RAM, decode speed is pulled toward the slowest tier from which weights are read, and a single bandwidth number is no longer meaningful. The methodology page lists offload as not modelled. Fit still indicates whether the model loads, but the tok/s figure assumes that everything is in fast memory.

Prefill is a different regime. Reading the prompt is compute-bound, so it is estimated separately: prefill tok/s = dense FP16 TFLOPS x 1e12 x MFU / (2 x active parameters), with MFU of 0.45 for NVIDIA, 0.22 for AMD, 0.18 for Intel and 0.20 for Apple, calibrated on llama.cpp pp512 runs with Q4_K_M weights. Time to first token is the number of prompt tokens divided by that rate. MoE benefits prefill as well, since FLOPs scale with active parameters, but the error bar is wider, about 30%, and wider still where no verified compute figure is available. A fast-decoding card can therefore still impose a long wait before the first token.

Context length is a further, less visible source of error. The decode formula accounts for the weights and nothing else, so it ignores the growing KV cache that attention also reads on every token. The cache itself is straightforward to size: 2 (K and V) x KV heads x head dim x layers x 2 bytes (FP16) per token. Some architectures keep a full cache on only some layers (linear attention, sliding window), and the VRAM code accounts for that. The speed estimate, however, does not decrease with context, so long-context numbers are optimistic. Quantized KV cache is also not modelled.

Quantization carries its own overhead. Model size is derived from effective bits per weight, so Q4_K_M files are not exactly 4 bits per weight, and measured sizes from Ollama or llama.cpp are preferred over the computed estimate whenever they exist. Batching, speculative decoding, multi-GPU interconnects and flash-attention variants are likewise outside the model.

Fit, briefly

The fit side is simpler. VRAM is weights (parameters x bits per weight / 8) plus KV cache plus about 0.3 GB of framework overhead. When architecture data is unavailable, the fallback is a flat 10% overhead. "Fits" is judged with room reserved for a 16K-token context, because a model that loads with a 2K cache is not usable for a real conversation.

The arithmetic for an individual model is printed on its page, for example gpt-oss-20b. Measurements from llama-bench that disagree with an estimate by more than the stated error are welcome, since this is how the constants are corrected.

Top comments (0)