DEV Community

Hotragn
Hotragn

Posted on

Fan-Out: Concurrent Tool Calls Under Shared-State Hazards

Kaggle Benchmarking Challenge Submission

DEV × Kaggle Benchmarking Challenge · #kagglechallenge

Author: Hotragn Pettugani · Kaggle

The problem

Agents increasingly fire many tool calls in one turn. Shared state — files, accounts, locks, lifecycles — makes order and aliasing matter. Most benchmarks score single-call correctness. Fan-Out scores safety and near-optimal scheduling when N calls share hazards.

What Fan-Out measures

  • Safety: no hazard violation (overwrite, double-spend, lifecycle race).
  • Safe-and-optimal (SAO): safe and uses the oracle-optimal number of rounds.
  • Speed given safe: distance to the oracle-optimal safe schedule.
  • Conditions include true concurrent fan-out (C0), sequential baselines, and a "lie" condition.

Oracle: conflict DAG + exhaustive interleaving check. No LLM judge.

Day-1 kill gate — PROCEED

Hardest rung: C0 × N=8 (7 items). Pre-registered kill: any frontier model ≥90% SAO → pivot.

| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| google/gemini-3.8-flash | 85.71 | 85.71 | 1.0 | 100.0 | 7.0 |
| google/gemini-3.7-flash | 85.71 | 71.43 | 0.9667 | 100.0 | 7.0 |
| google/gemini-3.5-flash | 85.71 | 71.43 | 1.0333 | 100.0 | 7.0 |
| anthropic/claude-sonnet-5@default | 85.71 | 57.14 | 0.9111 | 100.0 | 7.0 |
| anthropic/claude-opus-5@default | 71.43 | 57.14 | 0.96 | 85.71 | 7.0 |
| openai/gpt-5.4-mini-2026-03-17 | 28.57 | 28.57 | 1.0 | 100.0 | 7.0 |

Ceiling on this rung is 85.71%, under the 90% kill line → keep Fan-Out.

Task: fan-out-hard

Smoke ladder (easier mix)

Task: fan-out-schedule

| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| google/gemini-3.5-flash | 100.0 | 100.0 | 1.0 | 100.0 | 16.0 |
| google/gemini-3.7-flash | 100.0 | 93.75 | 0.9792 | 100.0 | 16.0 |
| google/gemini-3.8-flash | 100.0 | 87.5 | 0.9635 | 100.0 | 16.0 |
| anthropic/claude-sonnet-5@default | 100.0 | 75.0 | 0.9271 | 100.0 | 16.0 |
| openai/gpt-5.4-mini-2026-03-17 | 100.0 | 50.0 | 0.8333 | 100.0 | 16.0 |

Smoke saturates on safety; the hard rung is where separation shows up.

Swarm family (cross-agent conflicts)

Task: fan-out-swarm

| Model | Safe % | SAO % | Mean speed|safe | Format % | n |
|---|---:|---:|---:|---:|---:|
| openai/gpt-5.4-mini-2026-03-17 | 58.33 | 8.33 | 0.6429 | 100.0 | 12.0 |
| anthropic/claude-opus-5@default | 91.67 | 8.33 | 0.6833 | 100.0 | 12.0 |
| google/gemini-3.7-flash | 100.0 | 0.0 | 0.675 | 100.0 | 12.0 |
| google/gemini-3.8-flash | 100.0 | 0.0 | 0.675 | 100.0 | 12.0 |
| anthropic/claude-sonnet-5@default | 83.33 | 0.0 | 0.635 | 100.0 | 12.0 |

Strong models often stay safe but miss optimal multi-agent schedules — the complementary story to single-agent Fan-Out.

Why this is novel

Not another "pick the right tool" suite. Fan-Out is about same-turn concurrency safety under shared-state hazards. Closest prior work (PeakBench) covers data-flow/resource capacity and leaves state hazards out. Runtime products (AsyncFC, EffectFence, Restate) guard execution; they do not score the model.

How to reproduce

kaggle b t run fan-out-hard -m gemini-3.8-flash
kaggle b t download fan-out-hard -m gemini-3.8-flash -o out/
Enter fullscreen mode Exit fullscreen mode

Limitations

  • Day-1 item count on the hard rung is small (n=7); expanding the ladder is next.
  • Some frontier runs returned format_ok=0 (quota/API) and are excluded above.
  • k=1 for the gate; k=3 comes after expand.

What’s next

Grow N=12, more open models, held-out one-line-fix test, Pareto plots of safety vs speed.

kagglechallenge

Top comments (0)