Originally published on rohitraj.tech
PrismML shipped 1-bit and ternary builds of Qwen3.6-27B on July 14, 2026 — 5.9 GB for ternary, 3.9 GB for 1-bit, running at 163 tok/s on an RTX 5090 and 11 tok/s on an iPhone 17 Pro. Every writeup leads with "retains 95% of baseline." Nobody breaks out the row that matters: tool-calling drops 80.0 to 66.0 at 1-bit — degrading 4.6x worse than math. For a model sold on laptop-local agents, that is the whole story. Here is the variant decision table, the runnable commands, the KV-cache trap that makes 5.9 GB of weights need 13.7 GB of RAM, and how I would ship this in production.
Read the full version with code samples, diagrams, and architecture details: Bonsai 27B: A 27B Model on Your Phone — and the One Benchmark That Collapses (2026)
More engineering notes: rohitraj.tech/en/notes
Top comments (0)