DEV × Kaggle Benchmarking Challenge · #kagglechallenge
Author: Hotragn Pettugani · Kaggle
The problem
Agents increasingly fire many tool calls in one turn. Shared state — files, accounts, locks, lifecycles — makes order and aliasing matter. Most benchmarks score single-call correctness. Fan-Out scores safety and near-optimal scheduling when N calls share hazards.
What Fan-Out measures
- Safety: no hazard violation (overwrite, double-spend, lifecycle race).
- Safe-and-optimal (SAO): safe and uses the oracle-optimal number of rounds.
- Speed given safe: distance to the oracle-optimal safe schedule.
- Conditions include true concurrent fan-out (C0), sequential baselines, and a "lie" condition.
Oracle: conflict DAG + exhaustive interleaving check. No LLM judge.
Day-1 kill gate — PROCEED
Hardest rung: C0 × N=8 (7 items). Pre-registered kill: any frontier model ≥90% SAO → pivot.
| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| google/gemini-3.8-flash | 85.71 | 85.71 | 1.0 | 100.0 | 7.0 |
| google/gemini-3.7-flash | 85.71 | 71.43 | 0.9667 | 100.0 | 7.0 |
| google/gemini-3.5-flash | 85.71 | 71.43 | 1.0333 | 100.0 | 7.0 |
| anthropic/claude-sonnet-5@default | 85.71 | 57.14 | 0.9111 | 100.0 | 7.0 |
| anthropic/claude-opus-5@default | 71.43 | 57.14 | 0.96 | 85.71 | 7.0 |
| openai/gpt-5.4-mini-2026-03-17 | 28.57 | 28.57 | 1.0 | 100.0 | 7.0 |
Ceiling on this rung is 85.71%, under the 90% kill line → keep Fan-Out.
Task: fan-out-hard
Smoke ladder (easier mix)
Task: fan-out-schedule
| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| google/gemini-3.5-flash | 100.0 | 100.0 | 1.0 | 100.0 | 16.0 |
| google/gemini-3.7-flash | 100.0 | 93.75 | 0.9792 | 100.0 | 16.0 |
| google/gemini-3.8-flash | 100.0 | 87.5 | 0.9635 | 100.0 | 16.0 |
| anthropic/claude-sonnet-5@default | 100.0 | 75.0 | 0.9271 | 100.0 | 16.0 |
| openai/gpt-5.4-mini-2026-03-17 | 100.0 | 50.0 | 0.8333 | 100.0 | 16.0 |
Smoke saturates on safety; the hard rung is where separation shows up.
Swarm family (cross-agent conflicts)
Task: fan-out-swarm
| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| openai/gpt-5.4-mini-2026-03-17 | 58.33 | 8.33 | 0.6429 | 100.0 | 12.0 |
| anthropic/claude-opus-5@default | 91.67 | 8.33 | 0.6833 | 100.0 | 12.0 |
| google/gemini-3.7-flash | 100.0 | 0.0 | 0.675 | 100.0 | 12.0 |
| google/gemini-3.8-flash | 100.0 | 0.0 | 0.675 | 100.0 | 12.0 |
| anthropic/claude-sonnet-5@default | 83.33 | 0.0 | 0.635 | 100.0 | 12.0 |
Strong models often stay safe but miss optimal multi-agent schedules — the complementary story to single-agent Fan-Out.
Why this is novel
Not another "pick the right tool" suite. Fan-Out is about same-turn concurrency safety under shared-state hazards. Closest prior work (PeakBench) covers data-flow/resource capacity and leaves state hazards out. Runtime products (AsyncFC, EffectFence, Restate) guard execution; they do not score the model.
How to reproduce
kaggle b t run fan-out-hard -m gemini-3.8-flash
kaggle b t download fan-out-hard -m gemini-3.8-flash -o out/
Limitations
- Day-1 item count on the hard rung is small (n=7); expanding the ladder is next.
- Some frontier runs returned format_ok=0 (quota/API) and are excluded above.
- k=1 for the gate; k=3 comes after expand.
What’s next
Grow N=12, more open models, held-out one-line-fix test, Pareto plots of safety vs speed.
Top comments (0)