DEV Community

Jeffrey Turov
Jeffrey Turov

Posted on

A 3B model that beats a 7B: failure-driven orchestration on uncontaminated knowledge

How a fully sovereign QA stack — local Wikipedia index, distilled Qwen2.5-3B reader, and crutches built only from measured failures — went from 33% to 52% on facts no LLM can have memorized, and matched a zero-shot 7B more than twice its size.


TL;DR

Configuration Post-cutoff-150 accuracy
3B naked (no retrieval) 0.0% [0–2.5]
7B naked 0.0% (by construction — see below)
3B + sovereign chain (start, Sep 29) 33.3%
3B + sovereign chain (final, Oct 1) 52.0% [44.1–59.8]
7B zero-shot, same chain 50.0% [42.1–57.9]

Every number comes from a deterministic gate with Wilson 95% CIs. Every rollback is published below, not just the wins. The benchmark is public on Kaggle (post-cutoff-150).


1. Why "post-cutoff": most retrieval benchmarks measure memory, not retrieval

HotpotQA (2018) sits inside the training data of every 2024 model. We measured it: our naked 3B scores 30.5% on HotpotQA — from pure memorization. On such a benchmark, injecting retrieved evidence can only hurt: the "does retrieval help?" question is rigged from the start.

So we built post-cutoff-150: 150 questions generated from local Wikipedia pages whose lead paragraphs cite 2025–2026 facts, with answers that are exact spans of the source text. The contamination filter is deterministic and brutal: if the naked 3B can answer a question without context, the question is rejected. Final control: naked model = 0.0% [0–2.5]. Whatever the pipeline scores from there is net retrieval value — nothing else.

2. The sovereign stack

  • Index: 7.2M Wikipedia pages + 11.4M redirects, FTS5, 49 GB, fully local (crash-proof indexer, resumable to infinity).
  • Reader: Qwen2.5-3B-Instruct + LoRA stack trained inside the harness (SFT on noisy evidence, then DPO — the only technique that ever broke a plateau, twice).
  • Orchestrator: deterministic arms chained as a superset — anchor extraction → lead fetch → gated joint reading → deterministic bridge hop (n-grams certified against the 7.2M titles) → full-text search retry → Wikipedia tables arm (wikitext parsing with rowspan inheritance) → semantic lens → live full-text search.
  • Gate discipline: a challenger is PROMOTED only if it beats the champion on the same 150 questions, through the same matcher, with Wilson CIs reported. Otherwise ROLLBACK — and a rollback is a success of the system, not a failure to hide.

Two infrastructure rules that cost us bugs before they became rules: never evaluate a model in memory right after training (post-train state measures 0% on a perfectly healthy model — save, reload from disk, then gate), and all extraction LLM calls at temperature 0 (temp 0.1 anchoring injected ±4–6 points of run-to-run variance into every gate we ever ran).

3. The campaign: 33.3% → 52.0% in four days

Step Mechanism Score Verdict
Baseline pipeline anchors → lens → joint read 33.3% —
+ fullpool show the reader the whole pool, not just top sentences 34.7% kept (reserve)
+ FTS-retry full-text enrichment when reading fails 39.3% kept
+ entropy-gated routing accept a read only if token entropy ≤ 0.6 42.0% PROMOTE
+ SLOTS 2-slot "preuve → réponse" prompt with deterministic verification 43.3% PROMOTE
+ deterministic anchors fix, echo-retry, FR meta-strip — 46.0% kept (reserve)
+ FR→EN translation arm French meta-questions translated once, read in English 48.0% PROMOTE
+ Wikipedia tables arm wikitext tables → "header: cell" lines → lens 48.7% PROMOTE
+ semantic lens 12B judges sentence relevance instead of token overlap 52.0% PROMOTE

What did not work — measured, gated, published:

  • Self-consistency ×5: null. The literature assumes random errors that voting smooths; ours are systematic (wrong page → same wrong answer 5 times).
  • Best-of-N by confidence, inter-path consensus: null, same reason — independent paths converge to the same errors.
  • IRCoT (academic interleaved retrieval-CoT): 10% naive, 25% hybrid — our sequential pipeline beats it by 17+ points. For a 3B, anchoring by exact titles beats conversational retrieval.
  • Query planner (typed decomposition): rollback both directions — deterministic decomposition doesn't convert hard failures; they're hard for rules too.
  • SpanLift (generation-free QA: enumerate spans, judge by evidence lift): 1.3% — catastrophic. The candidate enumerator produced sentence fragments, and the scorer missed even when the truth was a candidate. Paradigm abandoned on 3B.
  • GRPO (both reader and GSM8K): gradient desert is structural (frac_reward_zero_std = 0.8). We fixed the mechanism (dense rewards → 0.2, reward flowing) — and transfer to the target distribution was still null. Two independent adapters scored identically to the champion.
  • Deep plaintext retrieval (full pages beyond the lead) and live Wikipedia full-text search: both net zero — every win overlapped an existing arm's rescue.

4. The capstone: 3B trained + orchestrated vs 7B zero-shot

Same chain, same questions, same matcher, same deterministic gates. The only variable swapped: the reader (Qwen2.5-7B-Instruct, 4-bit, no LoRA, no tuning).

Arm wins 3B (78/150 = 52.0%) 7B (75/150 = 50.0%)
Primary reader arm 50 66
Crutch arms (semantic lens, tables, fallbacks) +28 +9

Raw capacity wins the primary arm (+16 for the 7B). But the crutches — each one built on a measured 3B failure mode (echo killed by SLOTS, entropy calibrated at 0.05–0.21 for correct vs 1.0+ for guessing, deterministic proof verification) — add 28 points to the 3B and only 9 to the 7B. The 7B has no echo to fix and its entropy sits near 0.002, making the calibrated gate useless. The crutches don't transfer because they were never generic — they were prosthetics fitted to a specific patient's limp.

Honest caveats: the confidence intervals overlap (we claim "ties or beats", not "beats"); the 7B is 4-bit quantized and zero-shot — a tuned 7B would do better, and that's a fine next experiment.

5. What we actually learned (the publishable part)

  1. Contamination decides what a benchmark measures. On memorized knowledge, retrieval adds nothing; on post-cutoff knowledge, a sovereign pipeline turns 0% into 52%.
  2. Small-model errors are systematic, not random. Everything built on smoothing randomness (voting, consensus, best-of-n) failed four independent times. Everything built on diagnosing the systematic class worked.
  3. Entropy is a calibrated confidence signal nobody was reading. Correct answers: 0.05–0.21. Echoes/guesses: 1.0–2.27. One threshold, zero training, +2 points.
  4. SLOTS = externalized working memory. Force the model to first copy the proof sentence, then copy the answer from inside the proof, and verify both deterministically at receipt. The meta-question echo class (24 questions) died (3 left).
  5. Context quality beats context coverage. The semantic lens won 8 net questions — 4× its coverage diagnosis — because a relevance-filtered context reads better, not just finds more. Measure arms in-chain, not by detection rate.
  6. Deterministic gates catch the bugs that fake results. In-memory post-training evals, a urllib import swallowed by try/except (the "reader" we benchmarked for days was the fallback), a wedged CUDA process serving empty answers, ±6-point anchoring noise — every one of these would have become a false claim without the discipline.
  7. DPO is the only training move that ever broke a plateau (GSM8K 60→67, reader 42→45, twice confirmed). SFT saturates; GRPO's gradient desert is structural.

6. Reproducibility

  • Benchmark: post-cutoff-150, public on Kaggle, with the contamination filter (naked = 0% verified).
  • All scores from deterministic gates (Wilson 95% CI), same matcher everywhere (a lower bound, consistently applied — comparisons valid).
  • 6 promotions, 6 rollbacks — all archived with per-question outputs.
  • Hardware: one RTX 5090. No API calls to any frontier model for the answers — the 12B orchestration LLM and the 3B reader are both fully local.

7. Limitations

150 questions is honest but small; we report CIs and refuse to claim deltas inside the noise. The matcher (containment + token subset) undercounts morphological variants — scores are lower bounds. The 7B comparison is zero-shot/4-bit by design (we isolated the orchestration variable, not the best possible 7B). The residual 72 failures split between evidence absent from every source we can reach (obscure entities, join-queries) and pure 3B reading capacity — that ceiling is documented, not hidden.


Built by an autonomous harness that diagnoses its own failures, trains only through deterministic gates, and publishes its rollbacks. The rollbacks taught us more than the wins.

Top comments (0)