DEV Community

Jangwook Kim
Jangwook Kim

Posted on Originally published at effloow.com

vllm-mlx on Apple Silicon: Continuous Batching Works, but Unified Memory Sets the Real Limit

Why We Brought This Tool Into Our Lab

The appeal of Apple Silicon inference is easy to understand. A sufficiently large Mac can keep quantized models, application services, and development tooling on one quiet machine without a remote GPU endpoint. The harder question is whether that machine can serve several interactive users without turning every request into a serial queue.

M4 Max continuous batching speedups (5 concurrent requests)Llama-3.2-1B-Instruct-4bit (5 req) 2.05xLlama-3.2-3B-Instruct-4bit (5 req) 1.51xQwen3-0.6B-8bit (5 req) 3.39xQwen3-30B-A3B-4bit (5 req) 2.38xQwen2.5-1.5B-Instruct-4bit (5 req) 1.64x

Batching five concurrent requests yields 1.51x to 3.39x aggregate-throughput gains on the cited M4 Max models, so the multiplier depends on model, architecture, and memory bandwidth rather than being fixed.

That is the problem we evaluated with vllm-mlx. It exposes OpenAI- and Anthropic-compatible APIs over native MLX execution and adds the scheduling features missing from a basic mlx-lm generation loop: continuous batching, paged KV storage, prefix caching, request scheduling, and streaming.

Continuous batching is the important part. Instead of finishing request A before starting request B, the scheduler admits new requests into an active batch and advances multiple sequences together. Model weights can then be reused across rows in the same decode step. Aggregate throughput rises, although each client may receive tokens more slowly because it shares the machine.

The operational distinction matters:

  • Simple mode processes one request at a time and avoids batching overhead.
  • Continuous-batching mode combines active requests and optimizes aggregate throughput.
  • Paged caching stores KV state in fixed-size blocks and can share repeated prefixes.
  • Prefix caching helps only when prompts actually share reusable content.
  • None of these mechanisms creates additional physical memory.

Our evaluation hit an important boundary before model execution: the available isolated host was Linux without the Apple Silicon, Metal, and model weights needed for these tests. MLX scheduling, KV growth, and unified-memory exhaustion cannot be reproduced faithfully on that platform. We therefore did not fabricate a local M4 result or present a mocked scheduler as a real benchmark.

Instead, we reviewed the installation instructions and CLI options, then assessed throughput and configuration trade-offs using the Apple Silicon benchmark cases for vllm-mlx continuous batching, mlx-serve performance, and mlxcel scheduling. We did not execute an installation or CLI contract test in the sandbox. We used the throughput values below as benchmark reference points, not as measurements from this Linux host or a validated production capacity envelope.

For our vllm-mlx throughput evaluation, we used the benchmark case with an M4 Max, 128 GB of unified memory, and five concurrent requests. Aggregate throughput increased from 328.1 to 1,111.8 tokens per second for Qwen3-0.6B-8bit, from 98.1 to 233.3 tokens per second for Qwen3-30B-A3B-4bit, and from 196.9 to 322.2 tokens per second for Qwen2.5-1.5B-Instruct-4bit. Those are 3.39x, 2.38x, and 1.64x gains respectively.

That spread is the real story. “Continuous batching” is not a fixed multiplier. Gains depend on architecture, model size, memory bandwidth, batch width, KV traffic, prompt shape, and whether the implementation has a genuinely batched forward path.

Hands-On Walkthrough: Setup, Execution & Output

We pinned the walkthrough to a specific release rather than installing whichever version was latest. For this review, that reference was vllm-mlx==0.4.1.

On an Apple Silicon host, our deployment recipe would be:

python3 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install "vllm-mlx==0.4.1" openai

vllm-mlx model inspect \
  mlx-community/Llama-3.2-3B-Instruct-4bit

vllm-mlx model acquire \
  mlx-community/Llama-3.2-3B-Instruct-4bit \
  --target-dir ./models/llama-3b-4bit

vllm-mlx serve ./models/llama-3b-4bit \
  --served-model-name llama-3b \
  --host 127.0.0.1 \
  --port 8000 \
  --continuous-batching \
  --use-paged-cache \
  --cache-memory-mb 2048 \
  --max-num-seqs 8 \
  --stream-interval 1 \
  --enable-metrics
Enter fullscreen mode Exit fullscreen mode

We would not expose this listener directly to an untrusted network. The production command should additionally set --api-key, place the process behind a health-aware proxy, and apply an external concurrency limit below the server’s theoretical sequence ceiling.

The built-in benchmarker provides a practical first pass:

vllm-mlx bench-serve \
  --url http://127.0.0.1:8000 \
  --model llama-3b \
  --prompts short,medium,long \
  --concurrency 1,2,4,8 \
  --max-tokens 256 \
  --repetitions 5 \
  --scrape-metrics true \
  --format json \
  --output bench.json
Enter fullscreen mode Exit fullscreen mode

The following is an illustrative reporting schema we would use for automation, not a verified bench-serve output format or a measured result. We would check the actual benchmark output on Apple Silicon before implementing a parser:

{
  "run_metadata": {
    "model": "llama-3b",
    "concurrency": 4,
    "max_tokens": 256,
    "repetitions": 5,
    "result_kind": "simulated-example"
  },
  "metrics": {
    "requests_completed": 20,
    "requests_failed": 0,
    "aggregate_output_tokens_per_second": null,
    "median_ttft_ms": null,
    "p95_ttft_ms": null,
    "peak_metal_memory_bytes": null
  },
  "note": "Populate metric values only from the target Apple Silicon host."
}
Enter fullscreen mode Exit fullscreen mode

For capacity testing, we would sweep three axes independently:

  1. Concurrency: 1, 2, 4, 8, and then higher only if memory headroom remains.
  2. Prompt length: approximately 512, 2,048, 8,192, and 16,384 tokens.
  3. Generated length: 128 and 512 tokens.

Mixing all three variables in one first run makes the result difficult to interpret. Long prompts stress prefill and KV allocation; long outputs stress sustained KV growth; higher concurrency multiplies active sequence state.

We treat the default --max-num-seqs 256 as a scheduler ceiling, not a safe production batch size. A 256-sequence limit says what the scheduler may admit. It says nothing about whether a specific model, context length, and Mac can hold those sequences.

For streaming, --stream-interval 1 sends every token immediately. Raising it to 2–5 reduces delivery overhead but creates chunkier output. Values of 10 or more favor throughput over perceived smoothness. We would benchmark this separately from model throughput because HTTP chunk frequency can distort client-side latency without changing core decode speed.

Evaluation Limits and Production Risks

The first failure was environmental and absolute: no Apple Silicon meant no honest MLX performance run. MLX’s Metal execution and Apple’s unified-memory behavior cannot be inferred from CUDA, CPU emulation, or a fake cache implementation. We stopped that path rather than manufacturing numbers.

The second problem is that memory controls are easy to overestimate. --cache-memory-mb 2048 limits cache consumption; it is not a two-gigabyte limit for the entire process. Model weights, temporary prefill tensors, active KV state, tokenizer buffers, Metal allocations, the server, and macOS still compete for unified memory.

The default memory-aware cache fraction is 20% of available RAM. That is useful for containing prefix-cache growth, but it does not guarantee that the runtime will survive a large active batch. Likewise, --memory-budget-gb applies to registry-managed model-weight residency. It is explicitly not a total runtime-memory cap and does not guarantee prevention of a Metal or MLX out-of-memory failure.

That leaves us uncertain about how vllm-mlx handles unified-memory exhaustion. We found no verified basis for promising transparent CPU spill, automatic model eviction, or graceful continuation of active generations. Our production workaround would be preventive:

  • Reserve memory for macOS and colocated services.
  • Keep the model-weight footprint well below physical capacity.
  • Cap request prompt length and max_tokens.
  • Set a conservative application-level concurrency limit.
  • Reject or queue excess work before it reaches MLX.
  • Restart the worker after a fatal Metal allocation failure.
  • Run the API gateway outside the inference process.

Paged cache is valuable, but its memory savings depend on the workload. It reduces fragmentation and enables block sharing when requests have common prefixes. In our cache evaluation, we treated the 80%+ memory-saving figure as applying to shared-prefix KV storage for 10 or more concurrent users sharing the same system prompt, not to total process memory. We would not apply that estimate to unrelated prompts, unique RAG context, or independent coding sessions.

Prompt length also changes the result. The vllm-mlx M4 Max batching table does not disclose enough prompt and generation details to turn its speedups into a universal sizing rule. For mlxcel, we evaluated a more tightly specified benchmark case: on an M1 Ultra, meta-llama-3.1-8b-instruct-4bit, a 512-token prompt, and 128 generated tokens, four clients reached 107.9 aggregate tokens per second versus 56.8 for one batched client. Mean time to first token under four-client load was 189 ms, compared with 3,257 ms in the serial parallel=1 baseline.

That result shows why scheduling helps under contention, but it does not establish performance at 16K prompts or on an M4 Max.

Long prompts can also stall active decodes unless prefill is chunked. mlxcel’s scheduler makes this explicit with a default 2,048-token prefill chunk and a grant interval that lets decode work proceed between prompt chunks. vllm-mlx exposes a batching scheduler, but we would still test mixed workloads rather than assume a long ingestion request cannot disrupt short chats.

Batching support also depends on model architecture. Some implementations advertise batching even when a model family executes one single-sequence forward per row and merely evaluates those graphs together. mlxcel explicitly clamps unsupported SSM, hybrid, and mixed-cache families to one active batching row. Its native batched-forward list includes Llama 3, Llama 4, Qwen 3, Qwen 3.5, Gemma 3, Helium, Muse Glimmer, and Qwen3-MoE in the referenced implementation. We therefore need to verify model support for each family rather than infer it from a server-wide feature flag.

For more deployment reviews that separate API compatibility from actual serving behavior, browse our infrastructure tools collection.

Scale, Latency & Cost vs. Alternatives

We compared the three MLX-serving approaches around the behavior that matters in production rather than headline single-stream decode.

Option Default concurrency posture Verified reference behavior Memory-pressure behavior Best fit
vllm-mlx Continuous batching is opt-in; max-num-seqs defaults to 256 On M4 Max 128 GB, five-request throughput improved 1.51x to 3.39x across the five models listed above Cache can be capped, but that is not a total MLX memory limit OpenAI and Anthropic compatibility, broad features, shared Mac services
mlx-serve --max-concurrent 1 by default About 1.6x total throughput at four-way concurrency on dense models, with per-request latency cost KV quantization trades speed for 2x or 4x KV-memory savings Performance-focused local serving with explicit tuning
mlxcel Four active sequences by default On M1 Ultra and the specified 8B workload, four clients delivered 107.9 aggregate tok/s and 189 ms mean TTFT Paged-KV admission control is designed to return backpressure under KV-budget pressure; the budget guard is inert on dense decode Strong scheduler controls, bounded queues, and explicit admission policy
Direct mlx-lm loop Usually one request at a time No batching benefit Application must manage all queueing and recovery Single-user tools and offline jobs
Remote GPU endpoint Provider-dependent Scales beyond one Mac when capacity is purchased Memory management is delegated to the provider Bursty traffic, larger models, high availability, geographic distribution

mlx-serve exposes a direct memory-for-speed trade through KV quantization. In our tuning evaluation, we treated KV quantization as a memory-for-speed trade: about 10% lower decode throughput at 2K–4K context, and about 18% lower at 42K with fused packed reads automatically enabled from 8K, compared with the earlier 47% penalty. We budgeted approximately 2x KV-memory savings for 8-bit mode and 4x for 4-bit mode. That is attractive on a 16 GB Mac, but it is not free capacity.

Its sliding-window optimization also demonstrates why context length cannot be reduced to one number. On an M4 Max, Muse-Glimmer 30B 4-bit decode improved from 24.7 to 40.2 tokens per second at 16K context and from 8.6 to 21.0 at 64K after every attention read was trimmed to the model window. Those are release-specific measurements, not expected vllm-mlx results.

Cost break-even

For an illustrative capacity model, assume:

  • Mac acquisition cost: $4,000.
  • Amortization period: 36 months.
  • Average inference-system draw: 150 W.
  • Electricity: $0.20 per kWh.
  • Continuous availability: 720 hours per month.

Monthly hardware amortization is about $111. Electricity adds roughly $22, producing a baseline of approximately $133 per month before support, networking, backup capacity, and engineering time.

At an assumed remote accelerator cost of $1 per running hour, the raw infrastructure break-even is about 133 accelerator-hours per month. At $2 per hour, it is roughly 67 hours. These are planning assumptions, not current provider quotes.

The calculation favors the Mac only when it stays useful. A lightly used machine assigned to one small model can cost more per completed request than an autoscaled endpoint. A Mac already purchased for development has a different equation because its marginal capital cost may be close to zero.

Availability changes the result again. One Mac is one failure domain. A production service requiring maintenance without interruption needs at least two hosts, a load balancer, model synchronization, monitoring, and spare memory on both machines. Teams comparing one Mac against a managed endpoint often omit that redundancy cost.

If you need a workload-specific capacity and failover model, contact our architecture team rather than sizing from peak tokens per second alone.

Our Final Verdict: When to Deploy, When to Skip

vllm-mlx is credible for shared Apple Silicon inference. Continuous batching produces meaningful aggregate-throughput gains, and its OpenAI and Anthropic surfaces make migration easier than building a custom MLX scheduler. The M4 Max reference results show that the gain can be substantial, especially for smaller models.

Our reservation is operational, not conceptual. Unified memory is the hard boundary, and cache limits do not convert vllm-mlx into a memory-safe multi-tenant platform. Production safety still depends on request admission, context limits, monitoring, and worker recovery outside the server.

Deploy this if:

  • You already own Apple Silicon capacity with enough memory for weights, KV state, and operating-system headroom.
  • Several users share the same model and generate concurrently.
  • Your workloads reuse system prompts or conversation prefixes.
  • You need OpenAI or Anthropic API compatibility without a cloud dependency.
  • You can cap prompt length, output length, queue depth, and effective concurrency.
  • A single-region, local-network service is acceptable.
  • You can benchmark the exact model family rather than assuming all architectures batch equally.

Hold off or avoid it if:

  • You need guaranteed graceful recovery after unified-memory exhaustion.
  • Traffic is highly bursty and the Mac would sit idle most of the month.
  • You require multi-host scheduling, automatic failover, or elastic scaling out of the box.
  • Your workload depends on very long, unrelated prompts with little prefix reuse.
  • You are trying to run a model whose weights already consume nearly all physical memory.
  • You need a published service-level objective without first measuring mixed prompt lengths and generation sizes.
  • You intend to treat --max-num-seqs 256 as a safe batch size.

Our production starting point would be four active requests, a two-gigabyte cache cap, paged caching, strict API-side context limits, and at least 20% physical-memory headroom after model load. We would then raise concurrency only after recording aggregate throughput, per-request decode rate, median and p95 time to first token, queue wait, and peak Metal memory.

The bottom line is straightforward: vllm-mlx makes a Mac a legitimate multi-user inference server, not an elastic GPU cluster. Within that boundary, continuous batching is worth enabling. Outside it, especially when memory exhaustion or host failure must be invisible to clients, we would keep the Mac as an edge or development tier and route overflow to a managed backend.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to