DEV Community

Cover image for Dependency Bench: at what size do LLMs lose track of a CI pipeline?
Muhammad Owais Warsi
Muhammad Owais Warsi

Posted on

Dependency Bench: at what size do LLMs lose track of a CI pipeline?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Every engineer has stared at a red CI run and asked: why did deploy-prod get skipped when build-api only flaked
once?
Answering that means tracking dependencies, trigger rules, retries, timeouts and execution order across a
whole graph. AI agents are increasingly asked to do exactly this: debug pipelines, fix lockfiles, decide what to
rebuild. So I wanted to measure whether models can actually reason about dependency graphs, and at what size they
stop being able to.

Dependency Bench has 60 generated tasks: 5 families × 4 sizes (S → XL) × 3 instances.

Family The model must answer
Failure propagation (YAML) The final conclusion (success / failure / skipped) of every job in a CI workflow with needs, if: always() / failure(), continue-on-error, retries and timeouts
Failure propagation (prose) The same pipelines, described in shuffled plain English the way a teammate would explain them
Timing The finish minute of every job and the total duration, when only 2–4 runners are available
Version resolution One version per package satisfying every constraint, or UNSAT. "Take the latest" never works
Incremental rebuild Which Makefile targets rebuild after some files change, where some targets produce byte-identical output and stop the change from spreading

Sizes go from 10 to 200 jobs/targets and from 5 to 22 packages.

Why it's trustworthy. No LLM judge is involved. Every task comes from a seed, and the answer key comes from a
small reference simulator that implements the rules written in the prompt. Grading is exact: one wrong job anywhere
fails the task. I also record partial credit (share of jobs/targets correct) to see how close a failing answer
was. A perfect "oracle" model scores 100% through the same pipeline, which checks the grader end to end.

Here's a small rebuild task. Can you solve it?

out/http/schema: src/io/config.yaml
out/json/image:  out/http/schema
out/ui/test:     src/store/assets.json src/proto/config.yaml
out/metrics/lib: src/core/main.c out/ui/test
Enter fullscreen mode Exit fullscreen mode

Changed: src/io/config.yaml, src/store/assets.json. Identical output: out/http/schema.

(out/http/schema rebuilds, but its output is unchanged, so out/json/image does not rebuild.
out/ui/test and out/metrics/lib do.)

Models Tested

I ran the full suite locally through the OpenAI API at each model's default reasoning effort, with one tier of each
size so the curve has a top, a middle and a bottom:

  • gpt-5.6-sol: the strongest model I had access to
  • gpt-5.6-terra: the mid-tier model
  • gpt-5.4-mini: a small, fast model

Each task is a single user message: no system prompt, no tools, no code execution. The model has to reason, not
write a topological sort. On Kaggle, the same benchmark also runs on Claude models (Opus 5.5, Sonnet 5.5, Haiku 5.5)
next to these three.

Findings

Results viewer

Model Exactly right Partial credit S M L XL
gpt-5.6-sol 57/60 99% 15/15 15/15 14/15 13/15
gpt-5.6-terra 46/60 97% 15/15 13/15 12/15 6/15
gpt-5.4-mini 4/60 45% 2/15 1/15 0/15 1/15

1. Accuracy falls off with size, not with difficulty of the rules. terra is perfect on every family at size S, and
the rules don't change between S and XL. Only the graph gets bigger. Yet it drops to 6/15 at XL. The model knows the
rules; it loses track of them over a long chain.

2. "Almost right" is the typical failure, and in CI that's the dangerous kind. terra's failed 200-job answers
still got ~98% of jobs right. Its typical mistake is calling a job failure when it was actually skipped, mixing up
"this job broke" with "this job never ran because something upstream broke". That's exactly the kind of answer that
sounds convincing in a postmortem and is still wrong.

3. The same pipeline is harder in prose. On 200-job pipelines, terra scored 3/3 when the pipeline was YAML and
0/3 when the identical pipeline was described in English
. Structure is doing a lot of the model's work. Agents
that read incident descriptions, Slack threads or logs instead of config files should be expected to do worse.

4. Incremental rebuild was the only family that beat the strongest model. sol was perfect on failure propagation,
timing and version resolution at every size, but scored 2/3 on L and 1/3 on XL rebuilds. Its mistakes over-propagate:
it marks targets as rebuilt even when none of their inputs changed. In one case its own answer said a target's only
prerequisite was not rebuilt, yet listed the target as rebuilt anyway. "Identical output stops propagation" (the
idea behind restat/early cutoff in real build systems) seems to be the rule models apply least reliably.

5. Not reasoning is the cliff. gpt-5.4-mini answered in ~2 seconds with ~400 output tokens per task (sol: ~25 s,
~2,100 tokens) and got 4/60. On version resolution it declared UNSAT ("impossible") on 4 of the 8 puzzles that
did have a solution. That's the worst outcome for a dependency solver: giving up confidently.

What surprised me: how good the top model is. My first pilot was 5 deliberately tricky mid-size tasks, and sol
solved all of them. The interesting signal only appeared once I stopped hand-writing hard cases and started scaling the
same rules up
, which is also what happens in real monorepos.

What I'd measure next:

  • The same suite at higher reasoning effort, to see whether the size cliff moves or disappears.
  • An agentic track where the model can run the simulator or write code, to separate "can't reason" from "won't write the obvious script".
  • Real-world semantics: GitHub Actions' actual if: defaults, needs with matrix jobs, pip/npm resolvers.

My Benchmark

Built with Python. The task generator, reference simulator, exact grader and results viewer are all deterministic
from a seed, so the suite can be regenerated at any size to stay ahead of contamination.

Top comments (0)