My background spans ML accelerator compilers and OS virtualization, and I now work on accelerator fleet efficiency.
To explore how compilers, runtimes, and hardware interact in LLM serving, I built mini-vllm-rs, a single-node inference engine written entirely in Rust for macOS. It supports continuous batching, paged KV caching, prefix caching, and speculative decoding. CPU and Metal GPU workers can run concurrently, each handling complete requests.
The project and two articles document that learning process (with more to come):
- mini-vllm-rs: code, setup, and benchmarks
- Building an LLM Inference Engine with AI: The Code Was the Easy Part
- When Does Heterogeneous Inference Pay Off?
What building the engine taught me
To give you a glimpse of the problems I encountered, two requests generating 512 tokens each achieved an aggregate throughput of 16.8 output tokens/s when run sequentially on an Apple GPU. However, continuous batching reduced it to 8.7 tokens/s.
Why? Batching changed the matrix shape, making Candle, the framework I use for backend execution, switch from a GEMV kernel to a GEMM kernel. GEMM was slower for this small batch, so I added a GEMV fallback for small batches and benchmarked the crossover point to choose the switching threshold.
The first article follows how a simple generation loop grew into an engine with clear boundaries between scheduling, model execution, and request state. It examines how continuous batching required each request’s KV state to be managed independently of model execution, and how prefix caching complicated when that memory could be reused or released. Those boundaries later made adding concurrent CPU and GPU workers possible without a major redesign.
It also reflects on using AI to implement the engine, and the design reviews and experiments I needed to decide which changes to keep.
Why split work when a GPU can run the whole model?
In my project, I considered GPU prefill with CPU decode, or a GPU target model with a CPU drafter. Apple Silicon's shared memory looked helpful, but Candle's device-based memory model complicated cross-device state access. Synchronization and low request concurrency also made keeping both devices busy difficult. Hence, I implemented neither split.
Meanwhile, disaggregation is being explored at a much larger scale: NVIDIA has announced configurations that split inference work between Vera Rubin GPUs and Groq LPX, while AWS has announced prefill on Trainium paired with decode on Cerebras.
The second article asks when specialization can repay communication costs. In one simplified prefill/decode example, transferring 2 GiB of KV cache at 25 GB/s takes about 86 ms. Without compute–communication overlap, specialized accelerator decode must beat GPU decode by more than 0.67 ms/token for a 128-token output to repay that transfer. For a 512-token output, the threshold falls to 0.17 ms/token. The same accelerator could make a long answer faster but a short one slower.
Beyond this example, it compares prefill/decode, attention/FFN, and target/draft separation. It also asks why production might choose a slower path: more KV capacity, larger batches, and scaling each hardware pool independently can improve service capacity even when one request finishes later.
If you are coming to inference from another part of the stack, or mostly work with GPUs and want to understand the alternatives, you may be asking similar questions. I'm sharing the code, experiments, and articles as I learn, and I'd be interested to hear which trade-offs you've encountered.

Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.