Every person who has tried to make a 2B model do real work on their own machine knows the failure mode. The model loops on the same tool call, or declares victory with half the deliverables missing, or just goes silent in the middle of the task. The usual conclusion is "small models are toys".
I disagreed, so I built the thing I wanted to exist: Mingbird, an open-source agent harness for Windows + Ollama, designed for 2-9B models. Apache-2.0, GUI and CLI, no telemetry, no account, and out of the box nothing leaves your machine. One sentence on what it actually does: small models usually know what to do, they just cannot emit every step reliably or carry a task through to the end. Mingbird handles that layer, and a 2B model finishes real work.
That is the claim. The rest of this post is the evidence for it, because I would not believe me either.
The experiment
Four agent harnesses (Mingbird, goose, opencode, agent-mini, all stock, unpatched) × four open models (gemma4:e2b 2B, qwen3.5:4b, gemma4:12b, ornith-1.5:35b) × 18 real tasks = 288 cells. One Windows machine, one set of time budgets, temperature 0, thinking disabled, identical across all four arms. Tasks are real work: organize this folder, write a script and make it pass, multi-step long-horizon jobs. Scoring is deterministic and only reads the artifacts a run actually produced. It never reads the transcript, so a model that talks a good game but produces nothing scores nothing.
Results
| Model | Mingbird | goose | opencode | agent-mini |
|---|---|---|---|---|
| gemma4:e2b (2B) | 0.821 | 0.271 | 0.017 | 0.246 |
| qwen3.5:4b | 0.876 | 0.801 | 0.465 | 0.706 |
| gemma4:12b | 0.906 | 0.772 | 0.539 | 0.576 |
| ornith-1.5:35b | 0.941 | 0.679 | 0.896 | 0.092 |
Overall: 0.886 / 0.631 / 0.479 / 0.405.
Look at the shape, not the averages. At 2B, three of the four harnesses fall off a cliff and Mingbird does not: 0.821 is within 0.12 of its own 35B score. At 35B the field closes. Harness quality pays off exactly where the hardware is cheap, which is the tier most people actually own.
One model per tier conflates tier with family, so the entire four-harness matrix was re-run on a second small model from a different family: qwen3.5:2b, same code, same budgets, same scorer, 72 more cells, all published. 0.779 for Mingbird vs 0.239 / 0.096 / 0.017 for the others. Same cliff, different family. It is a harness property, not a lucky model pick.
The single cleanest data point
One task (WF-08), same 4B model, two builds of my own harness:
- Before one fix: 156 tool calls, zero files produced.
- After: 15 tool calls, three files.
- The bytes fed to the model were identical between the two runs.
Nothing about the model changed. What changed was how the harness handles loops, completion checks and tool calls.
What Mingbird does differently (3 of 10 mechanisms)
- Finish gate. Before accepting "done", the harness re-reads the original task and makes the model check its deliverables against it. Small models declare victory half-done; this was the single biggest win.
- Verification feedback. The harness runs the task's own tests and feeds the real error (file, line, message) back as a tool result, so the next step starts from what is actually broken instead of what the model believes is broken.
- Anti-loop. A repeated tool-call signature, or 15 turns with no output, marks the run as dead: escalating hints first, then a hard reset. This is the fix behind the 156-to-15 story above.
Plus one unglamorous mechanism: flat prefill. Tools load by category, and the factory prefill is pinned at exactly 797 tokens by a CI test that fails if it grows by one byte. A 2B model's context window is a scarce resource.
External check: tau2-bench
Self-built benchmarks deserve suspicion, so the same local qwen3.5:4b ran through τ²-bench (Sierra Research), all three domains, the same model in every arm. Retail: 0.763 vs 0.675 (the benchmark's native agent) and 0.588 (opencode). Airline: 0.740 / 0.740 / 0.500. Telecom: 1.000 / 0.930 / 0.991. First or tied-first in all three.
The most interesting result: frontier model, four harnesses
Point all four harnesses at one hosted frontier model (a local shim makes it look like an Ollama endpoint; no harness is modified) and run the same 18 tasks: 0.997 / 0.989 / 0.925 / 0.478. A defective harness throws away more than half of a frontier model. Meanwhile the three good ones land within 0.07 of each other: once the harness is good, the model is the bigger lever. Both directions are true at the same time, and the 0.017-to-0.271 range at 2B is what "broken harness" looks like on local models.
Nothing leaves your machine
The other half of the project. Inference runs on your own Ollama, and out of the box there is no cloud anything: no telemetry, no account, no update check, no analytics. Nothing in the code packages, snapshots or uploads your working directory, including its .git history. Voice input uses a bundled local STT model. The only outbound traffic is web search when the model chooses it (configurable) and the MCP servers you configure yourself, and the offline toggle keeps web tools from being assembled into the prompt at all and architecturally disables the (opt-in, off-by-default) cloud-model path. Verifiable with netstat.
That is what makes it usable for work you are not allowed to upload: internal source, client files, contracts, anything under NDA.
Fine print
Single machine, one trial per cell, 18 tasks, greedy decoding: treat the tiers as directional, not statistics. The published matrix ran on the v1.5.0 code base (frozen for the campaign), not the current v1.8.2 release tag. The per-mechanism ablation deltas are smaller than single-cell execution variance, so I do not claim per-mechanism numbers. Everything is in the repo, including the cells where I lose.
Try it
The whole matrix ran on an integrated-GPU box (Intel Arc B390, 32 GB RAM), so 2-9B is comfortably inside what any gaming GPU already does, and the 35B MoE still runs end to end at roughly 26 tok/s with 128K context. Apache-2.0, 461 tests.
- Releases (Windows installers EN/CN, Linux and macOS tarballs): https://github.com/Mingbird/Mingbird-agent/releases
- Repo (tasks, scoring code, 288 cells as CSV, per-cell reproduction steps, about 30 minutes per cell on the test machine): https://github.com/Mingbird/Mingbird-agent
If you have a laptop and an afternoon, the fastest way to check any claim here is to run one cell yourself. And if you once tried to make small models do real work and gave up, tell me in the comments what you asked it to do; I can probably tell you which of the ten mechanisms was missing.



Top comments (0)