Bonsai 2 27B: How Ternary Weights and a Hadamard Rotation Compress a 27B Model to 5.9 GB
A 27-billion-parameter model that fits in under 6 GB and retains 98.2% of its full-precision benchmark performance sounds like a contradiction. Bonsai 2 27B, released by PrismML on September 17, 2026, makes that claim concrete — and the mechanism behind it is worth understanding in detail.
The Core Idea: Weights as Three Values
Most quantization schemes compress model weights from 16-bit floats to 8-bit or 4-bit integers, still representing dozens or hundreds of distinct values. Ternary quantization goes further: every weight in the language model backbone is restricted to exactly one of three values — -1, 0, or +1.
A ternary value carries log₂(3) ≈ 1.585 bits of information, compared to 16 bits in FP16. That theoretical floor translates to roughly 1.72 bits per weight in practice once you account for the grouping scheme Bonsai 2 uses: one shared FP16 scale factor for every 128 weights (called "g128" grouping). The scale factor recovers the actual magnitude needed for computation, while the ternary values encode the sign and presence of each weight.
The result is a language model that occupies 5.9 GB in its densest GGUF format — down from approximately 54 GB for the FP16 version of the same Qwen3.8-27B base architecture. That is roughly a 9x reduction.
Why Ternary Quantization Usually Fails — and What Bonsai 2 Does Differently
Sub-4-bit quantization has a well-known failure mode: outlier weights. In a typical weight matrix, a small number of values have much larger magnitudes than the rest. When you force everything into three levels, those outliers dominate the group's scale factor, and the majority of weights get rounded to zero or ±1 in ways that destroy the model's ability to reason.
Bonsai 2 addresses this with a blockwise Hadamard rotation. Before ternary assignment, each weight matrix is transformed by an orthogonal rotation applied in 1024-element blocks using ±1 signs. The Hadamard transform spreads outlier weight energy across the entire block, producing a more uniform distribution that is far easier to represent with three levels.
This rotation is folded into the stored weights offline — no runtime cost on the weight side. The tradeoff is that inference requires a matching orthogonal transform applied to activations, so the model cannot be loaded with standard MLX or llama.cpp loaders. PrismML ships custom kernels for Apple Silicon (MLX) and NVIDIA (CUDA).
The ternary representation applies end-to-end across embeddings, attention projections, MLP projections, and the LM head — what the model card calls "no high-precision escape hatches." The vision tower (0.92 GB) stays in FP16.
Benchmark Results: Where the Compression Holds and Where It Doesn't
PrismML evaluated Bonsai 2 27B across 14 thinking-mode benchmarks using EvalScope and vLLM on H100 hardware. The headline number is 84.78 average, compared to 86.32 for the FP16 baseline — a retention rate of 98.2%.
Breaking that down by capability area:
- Math (AIME 2026, AIME 2025, GSM8K, MATH-500): 96.57 — within half a point of full precision
- Coding (HumanEval+, LiveCodeBench v6, MBPP+, BigCodeBench): 81.58 vs. 82.17 for FP16
- Instruction following (IFBench, IFEval): 82.66 — actually above the FP16 baseline of 81.25
- Knowledge and reasoning (MMLU-Redux, GPQA Diamond): 83.95 vs. 86.66 for FP16
- Agentic tool calling (τ²-bench, BFCLv3): 77.57 vs. 79.74 for FP16
Structured tasks like math and code generation are relatively robust, while open-ended knowledge retrieval and multi-step tool use show slightly larger gaps. The instruction-following improvement likely reflects the Qwen3.8-27B base model's strengths rather than a quantization artifact.
Compared to other low-bit alternatives at similar sizes, the results are striking. A standard IQ2_XXS build of the same base model — which averages 2.8 bits per weight and occupies 9.4 GB — scores only 72.59 on the same benchmark suite. Bonsai 2 achieves 84.78 at 1.72 bits per weight and 5.9 GB. The Hadamard rotation appears to be doing real work.
Storage Formats and the "Bit" Label Problem
One of the more useful aspects of PrismML's release is its transparency about how storage format affects the actual bit-per-weight count. The same underlying ternary weights can be packaged in three ways:
| Format | Bits/weight | Size |
|---|---|---|
| GGUF PTQ1_0 (dense trits) | 1.75 | 5.95 GB |
| GGUF PQ2_0 (2-bit slots) | 2.13 | 7.21 GB |
| MLX 2-bit (scale + bias per group) | 2.25 | 7.67 GB |
The MLX format is largest because its grouped low-bit container stores both a scale and a bias per group of 128 weights, where ternary only strictly needs one. The underlying values are identical across all three formats — the difference is container overhead, not quantization quality.
The choice between PTQ1_0 and PQ2_0 depends on hardware: PTQ1_0 wins on memory-bandwidth-limited GPUs (RTX 4090, Ada-generation cards), while PQ2_0 wins on compute-bound hardware (H100, A100, Blackwell) where cheaper unpacking arithmetic matters more.
Running a 27B Model on a Laptop
The practical implication of a 5.9 GB model is that hardware that previously couldn't hold a 27B model at all can now run it interactively. The FP16 version requires roughly 54 GB of memory — well beyond any consumer laptop. The ternary version fits comfortably on Apple Silicon unified memory:
- Apple M5 Max: ~47 tokens/second
- Apple M5 Pro: ~28.7 tokens/second
- Apple M4 Pro: ~18 tokens/second
On NVIDIA hardware, throughput reaches 91–130 tokens/second on RTX 4090 and 5090 cards depending on packing format. The model supports a 262K-token context window via the Qwen3.8-27B hybrid-attention backbone (~75% linear attention). PrismML also reports 0.714 mWh per token on an RTX 4090 — 40% more energy-efficient than an 8B model in full precision, which matters for always-on local agents.
What This Means for Local AI Deployment
The broader significance of Bonsai 2 27B is that the gap between "runs locally" and "is actually useful" is narrowing. Previous sub-4-bit models often sacrificed too much capability in coding, vision, or agentic tool use to be practical for real work. Bonsai 2's benchmark profile suggests that ternary quantization with proper rotation can preserve the capabilities that matter most for agent loops, coding assistants, and multimodal workflows.
The model is released under the Apache 2.0 license and available on Hugging Face in both MLX and GGUF formats. For practitioners evaluating local deployment options, the key question is whether the 1.8% average benchmark gap relative to FP16 is acceptable for their use case. For most applications — coding assistance, document analysis, private inference — it likely is. For tasks that depend heavily on broad knowledge retrieval, the gap in the knowledge-and-reasoning category (83.95 vs. 86.66) may warrant closer evaluation.
Conclusion
Bonsai 2 27B demonstrates that ternary quantization, when combined with a well-designed rotation scheme and end-to-end application across all weight matrices, can compress a frontier-class model to a fraction of its original size without collapsing its reasoning capabilities. The Hadamard rotation is the key technical ingredient — it solves the outlier problem that makes naive sub-2-bit quantization impractical. The result is a 27B model that runs on a laptop, consumes less energy than an 8B full-precision model, and retains 98.2% of its benchmark performance.
Top comments (0)