Originally published on rohitraj.tech
Qwen3.8-27B is the first Apache-2.0 model that scores 61.7 on SWE-bench Pro and still fits on one 24GB GPU. Here is the working-developer build: which of the 790 GGUF quants to actually download (with KL-divergence data), the llama-server flags that matter, wiring it into Qwen Code natively and Claude Code through a router, the cost math against a cloud agent subscription, and the context-window ceiling nobody puts in the headline.
Read the full version with code samples, diagrams, and architecture details: Qwen3.8-27B as Your Local Coding Agent: 24GB Setup, Quant Pick, and Claude Code Wiring (2026)
More engineering notes: rohitraj.tech/en/notes
Top comments (0)