OpenAI's 10,000-Agent Swarm: What Small Teams Can Copy
You have 300 contracts to review, 80 vendor quotes to compare, or a pile of research to summarize, and one long chat with one model isn't cutting it. OpenAI just published a case where about 10,000 agents worked one problem in parallel, and the headlines are about mathematics and a data-privacy fight. The useful part for a small team is the engineering pattern, and the part you should copy first is the verification, not the swarm size.
What OpenAI actually claimed, and what is still unverified
OpenAI says an internal system produced a proof that 3D Navier–Stokes dynamics can develop a singularity in finite time. It published a writeup and a Lean formalization, and the full details are on OpenAI's announcement page. Navier–Stokes is one of the seven Millennium Prize Problems the Clay Mathematics Institute listed in 2000.
"Solved" is not yet the settled word. The Clay Institute said the problem "has apparently been settled" and described its evaluation as deliberately unhurried. Under Clay's published rules, a proposed solution must appear in a Qualifying Outlet, at least two years must pass after publication, and it must gain general acceptance in the mathematics community. Clay does not accept direct submissions. OpenAI says it does not intend to claim the prize, which carries $1 million.
So the accurate framing is: a credible claim with a formal proof artifact, under review, not a confirmed prize. I did not read the proof itself, and I won't characterize its mathematics beyond what OpenAI states.
The system also isn't something you can use. OpenAI describes it only as an internal model significantly more capable than GPT-6 Astra, its most advanced publicly available model. OpenAI has described the model only as internal, and it is not publicly available.
The swarm by the numbers
The numbers below all come from OpenAI or from press reports of OpenAI's statements. Keep the Navier–Stokes run separate from the larger project.
| Metric | Value | Scope |
|---|---|---|
| Concurrent agents | On the order of 10,000 | Navier–Stokes run |
| Time to resolution | About 88 hours from first agents launched (Saturday, Sept 5) | Navier–Stokes run |
| Messages exchanged | About 2.7 million | Navier–Stokes run only |
| Tokens generated | Roughly 130 billion | Navier–Stokes run only |
| Messages, all problems | 4.9 million | Whole project |
| Output tokens, all problems | About 300 billion | Whole project |
| Proof formalization and checking | 17 more hours, done by GPT-6 Astra | After the result |
| Cost | "Millions of dollars," per OpenAI executives | Reported by Axios |
Some context on how it unfolded. After rumors on Tuesday, September 1 that two Millennium problems had been resolved, nearly 100 agents spent about 50 hours on a disproof of Euler regularity. OpenAI then moved resources to Navier–Stokes and seeded the agents with the Euler result. The agents had a cached copy of the internet and code execution, and they were split into groups that could talk within the group.
The only cost figure OpenAI executives gave is "millions of dollars."
What Noam Brown says the architecture is worth
The team's own answer is that the model mattered far more than the swarm. Noam Brown said he would not attribute even 10% of the credit to the multi-agent architecture and credits the underlying model's strength (MindStudio's writeup). The agents were not trained specifically for Navier–Stokes.
He also said the science at 10,000 agents doesn't exist yet, because ablations comparing 1,000 against 10,000 agents are too expensive to run (StartupHub.ai). Nobody knows how much of the result came from scale. His other point is the one to take to your own work: parallelism depends on the domain. Math and web research parallelize well. Writing a novel does not.
That gives you a practical rule before you build anything:
- Parallelizes well: work that splits into independent pieces with a checkable answer. Examples are extracting fields from 500 documents, checking 60 vendor claims against public sources, or screening 200 leads against a rubric.
- Doesn't parallelize: work where every piece depends on the last decision. Examples are a single narrative report with a consistent voice, or a negotiation strategy.
- Model quality first: if one strong model fails on a single item, 50 copies of a weaker model won't fix it. Test one item on the best model you can afford before you fan out.
The data question: what to check in your own tools
The controversy is a reminder to audit which of your AI tools train on your content. NYU mathematician Tristan Buckmaster said he and Anthropic researcher Levent Alpöge had put unpublished drafts into private Codex sessions. He asked whether OpenAI's model was trained on or had access to them. He said he was told the model did not look up user data but got no answer on training. He also said: "I am not accusing anyone of anything." (VentureBeat)
OpenAI's statement, as reported: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." (After the Update_) It denies that its researchers or agents accessed their specific user data, but says it can't entirely rule out an indirect connection (Axios).
The point for a small business doesn't depend on who is right. It depends on which plan you are on.
| Plan type | Default training use (per OpenAI's docs) |
|---|---|
| ChatGPT Enterprise, Business, Edu, and the API platform | Not used to train or improve models by default (OpenAI business data page) |
| Individual plans (ChatGPT and Codex) | May be used to train models unless you opt out; Codex has a separate "Include environments" setting (OpenAI Help Center) |
| API retention | Inputs and outputs may be retained up to 30 days for service and abuse detection, with exceptions (Enterprise privacy) |
Policies change, so check the current pages before you rely on this table. Here's a short audit you can run this week:
# Keep a plain-text register of every AI tool that touches client or proprietary material
cat > ai-data-register.csv <<'EOF'
tool,plan_type,training_default,opt_out_done,retention_notes,what_we_send
ChatGPT,individual,check-current-policy,no,check,draft contracts
Codex,individual,check-current-policy,no,check,internal repo
API pipeline,api,check-current-policy,n/a,check,invoices
EOF
Fill it in from each vendor's current documentation, not from memory. Anything sitting on an individual plan with unpublished client work should be moved to a business or API plan or have its opt-outs set. That's the cheap fix, and you can do it before the next project starts.
The pattern that transfers: fan-out, verify, cap
You can copy the shape of the swarm at 10 to 50 workers without copying its scale. Whether the 10,000-agent pattern transfers to small business pipelines isn't something any of the sources establish. This section is my own engineering judgment, based on Brown's remark that math and web research parallelize well. Treat it as a design to test, not a proven result.
The shape has four parts:
- Planner: splits the job into independent units with a clear "done" definition.
- Workers: run in parallel, each with one unit and only the tools it needs.
- Verifier: a separate call, or a deterministic check, that decides whether each result is acceptable.
- Budget cap: a hard limit on calls or tokens, so a bug can't burn through your bill overnight.
Here's a minimal version using the Anthropic Python SDK. Set the model to whichever current Claude model you've tested. Check the docs for current names and pricing.
import asyncio, json
from anthropic import AsyncAnthropic
client = AsyncAnthropic()
MODEL = "your-tested-claude-model" # check the current models page
MAX_CONCURRENCY = 10
MAX_CALLS = 400 # hard budget cap for the whole run
calls = 0
sem = asyncio.Semaphore(MAX_CONCURRENCY)
async def ask(prompt: str, max_tokens: int = 1000) -> str:
global calls
if calls >= MAX_CALLS:
raise RuntimeError("Call budget exhausted")
calls += 1
async with sem:
r = await client.messages.create(
model=MODEL,
max_tokens=max_tokens,
messages=[{"role": "user", "content": prompt}],
)
return r.content[0].text
async def extract(doc: str) -> dict:
out = await ask(
"Extract vendor, total, due_date, and payment_terms from this "
f"document. Reply with JSON only.\n\n{doc}"
)
return json.loads(out)
async def verify(doc: str, result: dict) -> bool:
# Independent check: the verifier sees the source AND the claim
verdict = await ask(
"Does every field below appear in the source text, exactly as "
"stated? Answer YES or NO only.\n\n"
f"SOURCE:\n{doc}\n\nFIELDS:\n{json.dumps(result)}",
max_tokens=5,
)
return verdict.strip().upper().startswith("YES")
async def process(doc: str) -> dict:
for attempt in range(2):
try:
result = await extract(doc)
except (json.JSONDecodeError, KeyError):
continue
if await verify(doc, result):
return {"status": "ok", "data": result}
return {"status": "needs_human", "data": None}
async def run(docs: list[str]):
return await asyncio.gather(*(process(d) for d in docs))
Three design choices in this code are drawn from the OpenAI story:
- Verification is separate from generation. OpenAI shipped a Lean formalization and had a model spend 17 more hours checking the proof. You can't do formal proofs on an invoice, but you can check that every extracted field literally appears in the source text.
-
Failures escalate instead of retrying forever.
needs_humanis a valid output. A swarm that never says "I don't know" will confidently hand you wrong answers. - The budget cap is a hard stop. With 130 billion tokens at stake, OpenAI presumably had cost controls. Yours should be a single integer at the top of the file.
To keep the cost estimate honest, work it as an example before you run anything. Say 200 documents each take two worker calls and one verify call, plus a retry on 10% of them. That's roughly 660 calls. Multiply by your average tokens per call and the current per-token price from the pricing page, and you have a ceiling before you spend a cent. This is illustrative arithmetic with round numbers, not a benchmark.
Where more agents stop helping
Adding workers speeds up independent tasks and does almost nothing for dependent ones. Reports on multi-agent "Ultra Mode" style setups describe roughly 2x speedup at higher cost, but the figures differ between secondary sources, so I won't quote them as fact. What you can rely on is your own measurement.
Run this experiment before scaling any pipeline:
experiment:
sample: 30 documents, hand-labeled by you
runs:
- name: single_agent_baseline
workers: 1
- name: five_workers_with_verifier
workers: 5
- name: ten_workers_with_verifier
workers: 10
record:
- accuracy_vs_hand_labels
- wall_clock_minutes
- total_calls
- needs_human_rate
decision_rule: >
Choose the cheapest config within 1 point of the best accuracy.
Brown's warning that nobody has run the 1,000 vs. 10,000 comparison applies to you in miniature. Most small teams never run the 1 vs. 5 vs. 10 comparison either, and then pay for parallelism that adds nothing. Thirty hand-labeled documents cost you an afternoon. Skipping the labels means you're guessing.
A few failure modes to watch for in your own runs:
- Correlated errors. If the worker and verifier use the same prompt style and model, they can share the same blind spot. Vary the verifier's instructions, or use deterministic checks like regex matches against the source.
- Context contamination. OpenAI let agents talk within groups. In a small pipeline, letting workers see each other's output often spreads one wrong assumption across all of them. Start with isolated workers and a single reducer.
- Silent truncation. Long documents that get cut off produce plausible-looking partial results. Log input length per call and flag anything near your context limit.
How BizFlowAI approaches this
I'm Lazar Milićević, a senior engineer, and I build AI automations for solopreneurs and small teams: working systems, no buzzwords. For a small team, the scale that matters is tens of workers, not thousands, and the effort belongs in the verifier, the budget cap, and the escalation path.
The jobs that fit best are the parallelizable ones: contract and invoice extraction, vendor and competitor research, lead screening against a rubric. If you have a pile of documents or research that one chat window can't handle, I'd recommend starting with the register of AI tools and data from above, since where your documents are sent matters as much as what the agents do with them.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Top comments (0)