Open-weight models are not a fixed number of months behind proprietary models. They move on several clocks at once.
Across 19 dated model snapshots from May 25 through September 6, 2026, I found open-weight leaders reaching an earlier proprietary benchmark frontier in 24 to 58 days on coding, intelligence, and agentic scores. But they did not erase the live gap. Proprietary leaders advanced during the same window.
That distinction changes the answer. Open weights are catching the last frontier quickly. Catching the moving frontier remains a different problem.
TL;DR
- In a same-methodology comparison, open-weight leaders reached within 1% of an earlier proprietary frontier in 28 days for coding and 58 days for aggregate intelligence; an open agentic leader exceeded a 24-day-old proprietary score.
- The live proprietary frontier still finished 2.7% to 5.4% ahead of the open leader on those same dimensions.
- External research points in the same direction: the aggregate lag has compressed from roughly a year in older benchmark history to about four months in early 2026, but no single lag applies across coding, agents, long context, or multimodal work.
- Benchmark catch-up is not product catch-up. Serving, tools, recovery loops, safety controls, and operations remain part of the system.
The question needs two clocks
Most catch-up claims compare a new open model with an older proprietary model. That is a valid diffusion measure: how long did it take an open release to reproduce a capability that was previously closed?
It is not the same as asking whether the best open model now matches the best proprietary model.
I used two clocks:
- Prior-frontier catch-up: the first date an explicitly classified open-weight leader reached at least 99% of an earlier proprietary leader's score.
- Live-frontier distance: the open leader's score relative to the proprietary leader available at the end of the comparison window.
I applied those clocks only inside methodology-compatible snapshots. The model dataset changed from methodology v4.1 to v4.2 in September, so I did not connect raw scores across that boundary. Open-weight status comes from a canonical, source-linked registry rather than a hosting-provider label. The analysis also fails if any top-25 model in a measured dimension lacks a verified classification.
The clean comparable window runs from June 21 through August 18.
The model names and scores below come from the dated snapshots behind the site's language model analysis.
| Dimension | Earlier proprietary frontier | Open-weight catch-up | Time | Open score versus live frontier |
|---|---|---|---|---|
| Intelligence | Claude Fable 5: 59.9 on Jun 21 | Kimi K3: 59.7 on Aug 18 | 58 days | 94.6% |
| Coding | Claude Fable 5: 76.5 on Jun 21 | Kimi K3: 76.2 on Jul 19 | 28 days | 97.3% |
| Agentic | Claude Opus 5: 55.3 on Jul 25 | Qwen3.8 2.4T A95B: 57.1 on Aug 18 | 24 days | 96.5% |
The gray baseline is the earlier proprietary score. Green is the open-weight leader by August 18. Orange is the live proprietary leader at the end of the same v4.1 window.
Coding closed fastest. Kimi K3 reached 99.6% of the June coding frontier in four weeks. Agentic performance moved faster still: Qwen3.8 exceeded a 24-day-old proprietary score, although the live proprietary leader had already moved another 3.7% higher. Aggregate intelligence took almost two months and retained the widest live gap.
This is the moving-frontier effect in data. A model can catch up and remain behind at the same time.
The September snapshot does not show universal parity
The September 6 snapshot uses methodology v4.2, so its raw scores should not be compared with the earlier series. It can still show the open-versus-proprietary distance within that single snapshot.
| Dimension | Open-weight leader | Proprietary leader | Open as share of proprietary |
|---|---|---|---|
| Intelligence | 50.2 | 56.8 | 88.4% |
| Coding | 76.2 | 81.6 | 93.4% |
| Agentic | 53.6 | 58.2 | 92.1% |
Math is absent from this table for a reason. The current source contains no usable math-index observations. Older rows would create the appearance of a current ranking from stale data, so the honest result is unavailable, not zero and not parity.
The same caution applies to hosted speed. The fastest explicitly classified open endpoint in the snapshot produced roughly one-fifth the tokens per second of the fastest proprietary endpoint. That measures a provider's serving stack, hardware allocation, batching, and load as much as it measures a model. It is an operational result, not an inherent property of downloadable weights.
The economic frontier tells a different story. Nine of the 14 intelligence-price Pareto points in the September snapshot are verified open-weight models; five are proprietary. Open weights can become rational deployment choices before they become the highest-scoring models. That is why capability ranking and economic ranking are separate questions.
The longer record says the gap is compressing
The broader research record supports a shrinking lag, but not one universal number.
Epoch AI's historical open-model study examined benchmark catch-up through September 2024. Individual observations ranged from 5 to 22 months, with a mean near 13 months. Its combined benchmark-and-training-compute estimate placed open models about 14 months behind.
A newer Epoch Capabilities Index analysis estimated an average open-versus-closed lag of about four months from January through May 2026. Requiring the open model to strictly exceed the earlier closed model's point estimate extended the estimate to six months.
Those results do not conflict with the 24–58 day windows in my snapshots. They answer different questions:
- Epoch measures an aggregate frontier across a longer history and multiple capability tests.
- My shorter result tracks the fastest explicitly classified open leader in each dimension during one active release window.
- A 99% threshold detects practical score proximity earlier than strict exceedance.
The historical direction is still clear. The diffusion window has compressed from roughly a year to months, and sometimes to weeks on a narrow benchmark. The remaining distance is jagged rather than uniform.
Stanford's 2026 AI Index illustrates that movement through human-preference scores. It reports the open-versus-closed Arena gap narrowing from 15.2% in May 2023 to 0.5% in August 2024, then reopening to 3.4% by March 2026. Catch-up did not end the race. A later proprietary release widened the gap again.
Each capability moves on a different clock
Coding is the fastest visible catch-up
Coding has structured tasks, executable outputs, and strong feedback loops. That makes it easier to generate training signals and verify improvement. It also makes benchmark proximity more credible than in subjective tasks.
But a coding score still does not establish application reliability. One 2026 study comparing benchmark rank with an end-to-end application task found that SWE-bench position did not predict which model produced the best working application in its experiment. Repository navigation, environment setup, tool use, and error recovery can reorder the models.
Agentic scores are close; agent systems are not interchangeable
Agent benchmarks depend on the harness around the model: tool descriptions, prompt templates, parsers, turn limits, state handling, and retry logic. Artificial Analysis gives agentic evaluations a large share of its intelligence methodology, but that still measures a configured evaluation system rather than a bare checkpoint.
The deployment requirements make the distinction concrete. The Qwen3.8 model card specifies serving engines, templates, and tool parsers. The Mistral Large 3 model card describes a reference configuration built around eight H100 or B200 GPUs. By contrast, the GPT-OSS-120B card targets a single 80 GB GPU.
Weight access gives control. It does not make every checkpoint equally simple to operate.
Long context and multimodality do not have a stable lag yet
Advertised context length is an input limit, not a measure of reliable reasoning across that input. Long-context tests vary by document structure, retrieval pattern, distractors, answer location, and output demands. I found no credible public estimate that converts those differences into one open-versus-proprietary catch-up duration.
Multimodal results are similarly benchmark-specific. The Stanford AI Index shows open models close on some image and video evaluations while retaining wider gaps on others. Combining those into a single time lag would imply a precision the evidence does not support.
Efficiency can arrive before capability parity
Open weights change the economics even when the benchmark leader remains proprietary. Builders can quantize, fine-tune, place inference near data, choose the serving stack, and trade throughput against quality.
Epoch's historical analysis found examples of open models matching older closed-model benchmark performance with materially less training compute. A separate analysis of inference price-performance trends found rapid annual improvement on open-weight Pareto frontiers. That advantage does not make self-hosting free: accelerators, engineering, capacity management, observability, and safety still belong in the total cost.
Catch-up is a vector, not a date
The useful model has three clocks:
| Clock | What catches up | What it misses |
|---|---|---|
| Benchmark diffusion | A checkpoint reproduces an older score | The frontier may have moved |
| System diffusion | Tools, serving, monitoring, and recovery become dependable | Product quality is hard to reduce to one index |
| Economic diffusion | The workload becomes cheaper or controllable enough to deploy | Lower cost does not guarantee frontier capability |
This explains why people can look at the same market and reach opposite conclusions. One sees near-parity coding scores. Another sees a proprietary agent product that works with less assembly. A third sees an open model that is cheaper, customizable, and sufficient for the workload. All three observations can be true.
The fair comparison is checkpoint against checkpoint or complete system against complete system. Comparing an open checkpoint with a finished proprietary agent product credits one side for its platform while charging the other for components that have not been assembled.
What the data still cannot answer
The snapshot record is useful but incomplete.
- The dataset contains 19 usable language snapshots over a little more than three months. That is enough to observe a release cycle, not enough to establish a permanent trend.
- The September registry classifies 408 of 643 catalog rows. The unclassified tail remains outside open-versus-proprietary calculations, while every top-25 model used by this analysis has verified status and a public evidence source.
- The September methodology change breaks raw-score continuity with the summer series.
- The current source has no usable math-index observations.
- The dataset does not measure production reliability, safety operations, data residency, customization effort, or the engineering cost of self-hosting.
The conclusion has to remain inside those boundaries.
So what
Do not wait for a declaration that open weights have universally caught up. It will never arrive, because capability does not move as one unit.
Evaluate the dimension your workload consumes. Coding may be within weeks of a prior frontier. Agentic benchmark scores may be close while system reliability remains far apart. Economic value may already favor open deployment even when the highest composite score remains proprietary.
The operating question is not, "Are open models at the frontier?" It is, "Has the open stack crossed the capability, reliability, and cost thresholds for this workload?"
The open thread is whether the live capability gap now trends toward zero, or whether proprietary labs keep reopening it by turning checkpoints into more complete systems faster than the open ecosystem can assemble them.

Top comments (0)