DEV Community

RC
RC

Posted on Originally published at randomchaos.us

The benchmark score is the number to trust least

Claude Opus 5.5 scores 58 on Artificial Analysis's Intelligence Index, more than double the median of 26 for comparable reasoning models. That is the figure that gets quoted. It also tells you the least about what the model costs to run.

The same evaluation run that produced the 58 recorded everything else. Getting through the Intelligence Index took Opus 5.5 some 260 million output tokens. The median model needed 81 million. That verbosity shows up on the bill: Artificial Analysis puts the average at $5.98 per task to run the index. Input is $4 per million tokens and output $20, both at the top of the range against medians of $2 and $10.

The composite score hides the number an operator actually budgets against. A model that reaches the same answers in a third of the tokens, and lands a few points lower on the index, can be the better buy for a given workload. The headline will not tell you that.

Latency is the other figure the score buries. Artificial Analysis reports a time to first token of 724.53 seconds against a 3.77-second median. Opus 5.5 is a reasoning model, so that window covers the extended-thinking phase before the first answer token appears; the site's end-to-end response time folds thinking time in for exactly this reason. Once it starts writing it is quick, 96 tokens per second against an average of 76. The wait is front-loaded into the thinking phase.

For an interactive tool that difference is the whole user experience. For a batch pipeline it may not matter at all. The Intelligence Index scores neither case; it scores correctness and leaves you to discover the rest.

None of this makes 58 a bad result. It places the model among the leaders on capability, with a 1M-token context window and text-plus-image input. The point is narrower: a single composite number is a ceiling on what the model can do, not a profile of what it does on your traffic. Treat it as a filter for the shortlist, then measure the candidates yourself, on your own prompts, counting output tokens, dollars per task and time to first useful token.

If you are pricing Opus 5.5 specifically, the blended rate Artificial Analysis quotes is $2.94 per million tokens at a 7:2:1 cache-hit/input/output mix, and the model is served by nine API providers, so the per-token cost is worth comparing across them before you commit.

Top comments (0)