Claude Sonnet 5.5 and GPT-6.1 Sol shipped a day apart at the exact same price: $2 per million input tokens, $10 per million output.
That makes the price sheet the least interesting part of the comparison.
I pointed both at a real open source repo (TinyDB) through a small harness and graded 132 agent runs automatically. The first round ended in a 30/30 tie, so I wrote harder tasks and ran them at two effort levels. Correctness stayed tied. Cost, speed, tool use, and code-review tradeoffs did not.
The mechanism is boring, which is why it bites. A price per token says nothing about how many tokens a model spends to finish a task. The tokenizers don't count the same text the same way, and agents differ in how many tool calls and turns they burn on the way to a passing result.
The consequence: if you pick a coding model off the pricing page, you are comparing the wrong unit. Measure cost per graded task on your own repo.
Full results, including the parts that flatter neither model: Same Price, Different Habits
Top comments (0)