DEV Community

Ashraf
Ashraf

Posted on

The Redis Guy Wrote a Local LLM Engine That Refuses to Be General. That's Why It Works.

Salvatore Sanfilippo (antirez, the guy who wrote Redis) shipped a local inference engine called DwarfStar 4, or ds4. It hit the Hacker News front page this week and the repo is sitting around 23k stars after about two weeks.

The pitch is one sentence from the README:

The code is self-contained and deliberately narrow, not a general GGUF runner: you need to use the GGUF files the project produces.

Everyone else is racing to support every model on earth. antirez went the other way. And the numbers say it was the right call.

What it actually is

A C inference engine for high-memory Mac (Metal), CUDA and ROCm machines. It targets a short list of models: DeepSeek V4 / V4.1 Flash, GLM 5.x, and Qwen 3.8 Flash Next. MIT licensed.

Not Ollama. Not llama.cpp. It does not load your random Q4_K_M from Hugging Face. You use the GGUFs the project builds, and in exchange model loading, prompt rendering, tool calls, KV state, the HTTP server and the coding-agent integration are all tested together, not as separate layers that happen to mostly work.

Redis energy: pick one workload, be ruthless about it.

The tricks that make it fit

Asymmetric quantization. MoE models route each token through a handful of experts, and those experts hold most of the weights. ds4 crushes the routed experts to ~2-bit and keeps the sensitive parts (shared experts, projections, routing) at higher precision. You lose far less quality than a uniform 2-bit squeeze would cost you, because you only compress the stuff that tolerates it.

Disk-resident everything that can be. Compressed KV caches and fast local SSDs make long contexts practical. Qwen 3.8 Flash's big n-gram tables (about 95 GiB) are read from disk rather than loaded into RAM, which is how a ~42 GiB weight set is viable on a 64 GB Mac.

Persistent KV cache. Long prefixes get saved to SSD and resumed by prompt hash. If you've ever watched an agent re-prefill 60k tokens of repo context for the tenth time, you know why this matters.

Numbers

From the project site (fast-changing beta, treat as a snapshot):

Machine Context Prefill (t/s) Generation (t/s)
M5 Max 128GB 2K 790.2 39.4
M5 Max 128GB 65K 398.5 27.6
DGX Spark 128GB 2K 825.8 18.1
DGX Spark 128GB 65K 823.0 13.8

Look at the 65K rows. Generation on the Mac drops from 39 to 27 t/s at long context. That's a usable agent loop on a laptop, not a demo. Independent write-ups report similar ballparks (roughly 35 t/s on an M5 Max, ~10 t/s for the huge PRO model on a 512 GB M3 Ultra), so the claims aren't just marketing.

Try it

git clone https://github.com/antirez/ds4
cd ds4 && ./download_model.sh ds4f-q2
make            # macOS / Metal
# make cuda-spark   # Linux / DGX Spark
./ds4                         # interactive CLI
./ds4-server --ctx 100000     # OpenAI + Anthropic style API
Enter fullscreen mode Exit fullscreen mode

The server exposes /v1/chat/completions, /v1/messages and /v1/responses, so anything that speaks OpenAI or Anthropic wire format can point at it. The README names Claude Code, Codex CLI, OpenCode and Pi as supported clients. There's also a native ./ds4-agent with /save, /list and /switch <sha> for sessions stored in ~/.ds4/kvcache.

Who this is for

  • You have 64 GB+ of unified memory or a DGX Spark / Strix Halo box. Below that you're in SSD-streaming, "it technically runs" territory.
  • You run long agent sessions and are tired of API bills and re-prefill latency.
  • Privacy or air-gap requirements that rule out sending code to a hosted API.

Who should skip it: anyone who wants to swap between forty models weekly. That's still Ollama's job, and Ollama is fine at it.

The caveats, because there are real ones

  • It's beta and moving fast. Expect breakage.
  • The README says plainly it's built with heavy help from AI coding agents, humans leading ideas, testing and debugging. I'd rather have that stated than hidden, but read the code before you trust it with anything sensitive.
  • One reviewer reports CPU inference can kernel-panic macOS. Stay on the GPU paths.
  • PRO model support is experimental, and the distributed mode (two Macs over Thunderbolt) has no encryption or auth. Don't put it on a network you don't own.
  • You're locked to the model list. If your favorite model isn't supported, you're waiting on antirez.

Why this matters beyond one repo

US search interest in "local LLM" reportedly grew more than tenfold between September 2025 and August 2026. The tooling answer so far has been broad compatibility layers. ds4 is a bet that the next jump comes from the opposite direction: one engine, one model family, every layer tuned to the others.

I think that bet is right. General runners will always be the safe default, but the 2x gains are going to come from people willing to say "no" to 95% of use cases. That's the same trade that made Redis fast.

Clone it, run the Q2 model, and point your coding agent at ./ds4-server. Then tell me in the comments what your tokens per second look like.

Sources: ds4 repo, DwarfStar site, Frate Pietro's write-up, Hacker News front page.

Top comments (0)