Cross-post of Insights #7 — canonical: https://sheikhwasim.com/insights/cheap-twin-escalate-when-needed/
Most agent stacks send every request to the flagship model "just in case."
That looks careful. It is usually waste — and it hides when the cheap path was already good enough. Latency climbs. The bill spikes. Nobody can answer: which tickets actually needed the heavyweight brain?
I call the missing control a Cheap Twin: run the same input through a small model first. Escalate to the expensive brain only when the twin disagrees with a short checklist, or fails a smoke assert. Everything else ships on the twin — with a receipt that names which model answered.
Receipt Gate makes each write legible. Cheap Twin makes each brain choice legible — and defaults the happy path to the model you can afford to run all day.
Quick use case
Situation. A commercial construction ops team shipped an "RFI triage agent." It scored inbound requests for information: trade, urgency, and suggested owner. Leadership wanted accuracy, so every ticket went to the flagship model. Staging looked sharp.
What broke. Volume hit production. Latency climbed into the awkward zone where PMs refreshed and re-submitted. The model bill jumped without a matching jump in "hard" tickets. Nobody could say which RFIs needed the heavyweight path. Logs showed tokens and a model name — not why that model was chosen. There was no choice. There was a default.
The fix (Cheap Twin).
- Run the twin first — small model proposes trade / urgency / owner on every RFI
- Gate escalate on a short checklist — missing deadline, multi-trade conflict, owner-facing language, dollar threshold, unknown vendor
- Fail closed on smoke asserts — if the twin output fails a cheap structural check, escalate (or hold) instead of shipping mush
- Escalate with the twin's draft — flagship gets the same input plus the twin proposal and which checklist line fired
- Receipt the brain every ticket logs twin vs flagship, checklist hits, and cost — skim-worthy, not a token dump
Same tools. Same schema. Most traffic never left the twin. When it did, the receipt said why.
Cheap Twins aren't about starving the agent of intelligence. They're about refusing the expensive path when you have no signal you need it.
Flagship-by-default is a missing control
A flagship call answers: What would the best model say?
A Cheap Twin answers: Did we need the best model for this ticket — and can we prove it?
Those are different questions. Most stacks fund the first and call the bill "the cost of quality." Quality without a routing receipt is hope with a larger context window: you cannot tune the checklist, Golden-Trace the cheap path, or tell finance which spike was complexity vs lazy routing.
The Cheap Twin stack (five layers)
Build it like a control on the model path, not a discount coupon in the prompt.
1. Twin-first contract
Every eligible request hits the small model before the flagship. The twin returns a structured proposal in the same schema the flagship would use — not a chatty essay the orchestrator reinterprets.
If the twin cannot speak the production schema, it is not a twin. It is a different product.
2. Escalate checklist (five lines, not fifty)
Name the concrete reasons the cheap path is not enough. Keep the list short enough to argue in standup:
- Missing or ambiguous deadline
- Multi-trade / multi-system conflict
- Owner-facing or customer-facing language
- Dollar / risk threshold crossed
- Unknown vendor, SKU, or entity
Vague scores ("confidence < 0.7") belong behind a named line — or they become a knob nobody trusts.
3. Smoke assert before ship
Even on the twin path, run a cheap structural assert before anything leaves: required fields present, enums valid, no empty owner, no "see above" mush. Fail → escalate or hold. Do not ship a polite brick because the small model was fast.
Pair with your Smoke Evals pack so the same traps block CI and live routing.
4. Escalate with context, not from scratch
When a checklist line fires, the flagship should see the twin's proposal and which line triggered the escalate. That keeps the expensive call short and disagreement inspectable. Blind re-runs that ignore the twin waste the money you just decided to spend.
5. Brain receipt in the run record
Every ticket leaves a skim-worthy receipt: twin model id, flagship model id (or "not called"), checklist hits, assert pass/fail, latency, and cost. Mushy "used GPT" fails the gate — same rule as Receipt Gate.
Freeze a twin-happy path and a forced-escalate path as Golden Traces. Prompt and model bumps must not silently flip the routing mix.
How this maps to what you already have
- 5-layer agent stack — Cheap Twin lives in Models + Guardrails; Traces carry the brain receipt
- Smoke evals — "twin ships when checklist clear" and "checklist hit → escalate"
- Golden Trace — freeze one twin-only success and one escalate success; regress both
- Receipt Gate — model choice deserves a plain-language receipt before the next write
- Permission Envelope — escalate widens which brain may propose, not which tools
- Ship Gate — prove routing + receipts in canary, not that the demo always hits flagship
Architecture without a twin is just a flagship with a louder invoice.
The test
Ask one question before the next "always use the best model" ship:
For a normal ticket, does a small model propose first — and can a teammate skim a receipt that says whether you escalated, which checklist line fired, and what it cost?
If the answer is no, you are still buying flagship by default.
Closing
I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.
Follow for practical AI architecture that ships.
Connect on LinkedIn: Wasim Sheikh · Site: sheikhwasim.com · Notes: Practical AI Notes · X: @anciwasim
Top comments (0)