Karpathy's llm-council showed a lot of people something quietly important: the
most useful part of asking five models a question is not reading five answers. It
is watching them review each other. His weekend project runs three stages — every
model answers, every model ranks the anonymised answers of the others, a chair
model synthesises — and people keep finding that the ranking stage surfaces things
the first stage never said.
We have been building in the same direction (AI Group Call
— disclosure: our product; it runs each seat on a model you pick from the major
labs), and the pattern holds up in a different format: a live voice call where each participant
hears the whole conversation before its turn. Three things change when models stop
answering in isolation:
Positions move. In independent answers, nothing ever changes its mind. In a
shared conversation you see it happen: model B starts certain, hears model A's
constraint it hadn't considered, and walks back its own suggestion. That retraction
is the highest-signal moment in the whole session — it tells you which consideration
actually mattered, something five parallel chats never show you because they never
had to disagree.
Errors get caught by the room, not by you. One model hallucinating an API flag
tends to get corrected by another seat, roughly the way a colleague says "that
wasn't in the docs last I checked." It is not a substitute for verification, but it
beats you being the only reviewer of five confident paragraphs.
The format forces brevity. Spoken turns are short. Counter-intuitively, that is
a feature for decisions: you get "no, because X" instead of four thousand words of
"it depends." For long-form analysis, written councils still win — a voice room is
a debate club, not a research library.
Making the argument actually produce a decision
Running the room is a skill. What we have seen work across hundreds of sessions:
- Name a concrete goal, not a topic. "Should our solo app charge $9 or $19 a month?" gets a room moving. "Pricing" gets five essays.
- Give seats conflicting jobs. A champion, a CFO-shaped skeptic, a customer voice. Rooms where everyone is "helpful assistant" converge on mush.
- Make each seat commit before the discussion. Ask for one-line answers up front; the debate is then about why the answers differ, which is the whole point.
- Say the disagreements out loud. When two seats contradict each other, ask both to respond in one sentence. That is where the actual trade-off lives.
- End with the chair. Ask the most senior seat to summarise: what changed, what is still unresolved, what to do next. If you keep transcripts (ours saves them with every call), paste the summary where you work.
What this is not
Not a benchmark — confident is not correct, and a model agreeing with you is not
evidence. Not a code reviewer; spoken turns are too short for long proofs. Not a
legal or financial board. The niche where a live multi-model room genuinely beats
both single-model chat and multi-tab comparison is decisions with trade-offs:
architecture calls, pricing questions, "ship or polish" arguments, rehearsing a
pitch against investors who interrupt.
Whatever tool you reach for, the takeaway is architectural, not product-specific:
the unit of AI leverage is moving from the answer to the argument between answers.
Build your workflow around getting that argument — cross-review, ranked dissent,
or just a room where the models can hear each other — and the model you pick
matters a lot less than it feels like it should.
Top comments (0)