DEV Community

Cover image for What is disaggregated prefill and decode in LLM inference?
DigitalOcean for DigitalOcean

Posted on

What is disaggregated prefill and decode in LLM inference?

What is disaggregated prefill and decode in LLM inference?

Prefill is compute-bound, decode is memory-bound. Learn why colocating them hurts TTFT and ITL, and when disaggregated serving pays off.

Every LLM request runs through two phases with effectively opposite hardware profiles. Understanding the split explains most of the latency and cost behavior you see in production, and why disaggregated prefill and decode has become a standard serving pattern.

Key takeaways:

  • Prefill is compute-bound and sets TTFT; decode is memory-bandwidth-bound and sets ITL/TPOT.
  • Running both on the same GPUs creates prefill/decode interference, visible as TTFT spikes and ITL jitter.
  • Chunked prefill and continuous batching reduce interference; disaggregating the phases onto separate pools removes it, at the cost of a KV cache transfer.
  • Disaggregation is worth it for long prompts, high concurrency, and strict streaming SLOs. A managed option like DigitalOcean Serverless Inference gives you it without the setup; for small self-hosted deployments, a well-tuned single pool on a GPU Droplet running vLLM is often enough.

The two phases of LLM inference

Prefill processes the entire prompt in one forward pass and builds the KV cache, the attention state every later token reads from. It is compute-bound: arithmetic intensity is high and the GPU spends its time doing math, not waiting on memory. Prefill governs time to first token (TTFT), and its attention cost grows quadratically with prompt length.

Decode generates output one token at a time. Each step reloads the model weights and the growing KV cache from HBM to produce a single token, so it is memory-bandwidth-bound. Decode governs inter-token latency (ITL) and time per output token (TPOT). A GPU with higher memory bandwidth streams tokens faster even if its FLOPs are unchanged.

Why colocating prefill and decode causes latency problems

Most serving engines batch prefill and decode on the same GPUs. That is where the interference shows up: a long prompt's prefill can monopolize compute while dozens of in-flight decode requests stall, and users see ITL jitter or a TTFT spike. The two phases also want different batch sizes and parallelism strategies, but colocated they must share one configuration.

Chunked prefill and continuous batching soften this by interleaving prompt chunks with decode steps, which is why our inference trilemma guide recommends them for latency-sensitive deployments.

What disaggregation changes

Disaggregated serving, formalized in the DistServe paper, runs prefill and decode on separate GPU pools. A prefill worker computes the KV cache, ships it over a fast interconnect (RDMA, RoCE, or a transfer library like NIXL or LMCache) to a decode worker, and the decode pool streams tokens without prefill interference. Because each pool is sized and tuned independently, you can hit TTFT and ITL targets at the same time rather than trading one for the other. The cost is added complexity plus a KV transfer step; that transfer can land on second-token latency rather than TTFT depending on how orchestration is done.

This is the architecture DigitalOcean already runs in production: Serverless Inference uses disaggregated serving over a RoCE network, paired with prefix-aware routing so shared system prompts skip redundant prefill entirely.

References & further reading

Top comments (0)