DEV Community

Aamer Mihaysi
Aamer Mihaysi

Posted on

The AI time-horizon doubling rate is a curve-fitting artifact

Last month I sat in a planning call where someone said the line out loud — AI time horizons double every few months — and then used it to size a team. I've said that line myself. I've put it on slides. After a week of rebuilding my own eval harness, I think it's mostly an artifact of how we fit the curve, not a property of the models.

What the plot actually is

Strip the framing and it's simple. Take a task suite where every task carries a human completion time: 2 minutes, 15 minutes, 4 hours. Run the model on all of it. Bucket by human time, measure success rate per bucket, find the point where success crosses 50%. That crossing point gets called the model's "time horizon". Do that across a series of models, fit an exponential to the crossing points over calendar time, and out comes a doubling rate.

The trouble is the shape of the thing you're thresholding. On most of these suites, success between roughly 2 and 30 minutes of human work is nearly flat. A model that clears 70% of 5-minute tasks clears maybe 65% of 20-minute tasks. Nothing much happens. Then somewhere past that, the curve falls off a cliff, and the 50% crossing sits inside the cliff.

So the exponential fit is driven by the cliff. The flat region contributes almost nothing to the slope; the steep region contributes almost everything. What you get out is a number describing how fast a threshold moves through the part of the curve where the model was already competent — and it gets quoted as if it described the part where the model isn't. That's the myth. The doubling rate reads like a rate of capability gain. It's closer to a rate of threshold drift through the easy region.

Why the flat region exists

Long-horizon tasks mostly fail for boring reasons. A tool call returns garbage. A file path is wrong. The agent loses the thread after a context compaction and re-does work it already did. Those failures are roughly constant per step, so success decays multiplicatively with step count — which, plotted against log human-time, gives you a gentle slope early and a cliff later.

That mechanism matters because it's not the same as "the model can't reason for 20 minutes." It's "the harness leaks a little on every step." I've fixed more long-horizon failures this year by adding a schema validator on tool output than by swapping models. If your cliff is a harness artifact, your doubling rate is measuring your harness.

Three statistical problems

Task counts per bucket are small. A 50% crossing estimated from a couple dozen tasks has a confidence interval that can span hours of human time. Move two tasks between buckets and the crossing moves. I watched my own crossing point jump by roughly a factor of two after I changed how I scored partial credit — same model, same weights, same afternoon.

The functional form is a choice. Exponential, logistic, power law — inside the window where you have data, all three fit about equally well. Outside it, they diverge fast. The doubling rate is a parameter of a model you picked, not a measurement of the world. Fit a logistic in log-time and the headline changes while the data doesn't.

The crossing moves for non-capability reasons. Retry budgets, tool availability, grader strictness, whether the suite got easier between releases. Every one of those shifts the threshold. None of them is the model getting smarter.

The honest reading

A conversion curve. Human-minutes to AI-difficulty. It answers a question I actually have: of the work in front of me, what fraction is under the line today? That's a capacity-planning input. It is not a forecast, and it does not come with a date.

I've started reporting it that way. My own task log — the stuff I actually hand to agents, not a benchmark — buckets by human time, and I plot success against it. The median task in my log is a few minutes of human work, which means most of my backlog sits in the flat region where the curve tells me almost nothing about tomorrow. When someone asks "when will the agent do X", I answer with two things: where X sits on the curve today, and how much the curve moved since last quarter. Then I show the interval. The interval is usually embarrassing, which is the point.

What I'd want next

Per-domain curves instead of one aggregate — coding and research and ops have different cliffs, and averaging them produces a number nobody can act on. Published task counts per bucket. Confidence intervals on every crossing point. And an explicit refusal to extrapolate past the last measured bucket.

That last one is the hard one, because the extrapolation is the part everyone quotes. A flat-looking segment plus a steep segment plus a fitted exponential is a machine for generating confident dates out of thin data, and it's very good at it.

None of this means progress isn't real. It means the headline number is measuring the easy part of the curve and calling it the whole thing. If you're planning headcount off it, you're planning off the flattest data you have.

I haven't reproduced their fit myself — I read the paper, I didn't rerun their pipeline, so treat my reading as a reading. It's here if you want to check it against theirs: https://arxiv.org/abs/2610.12466v1

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

The multiplicative step decay point matches what I see in my own terminal every week. A 40-step task almost never fails because step 38 required deeper reasoning. It dies because every tool call has a low single-digit chance of returning an unhandled error, dropping an environment variable, or stumbling over a path that got mangled in a context compaction.

If step reliability is 97 percent, a 40-step loop is below 30 percent completion before the model even makes a hard semantic choice. Fixing the schema boundary, asserting postconditions on tool exits, and checkpointing filesystem state shifts that cliff further to the right than any prompt tweak I have tried. Conflating harness reliability gains with model capability doubling is how planning teams end up sizing roadmaps around statistical noise.

Collapse
 
devantibot profile image
DEV ANTIBOT •

You need to verify your account .
Link is in the profile.