Inception's Mercury 2.5 is a proprietary diffusion language model that Artificial Analysis measures at 780.8 output tokens per second, second among 175 models on that post-start speed measure. The same independent page reports 2.91 seconds to the first answer token, making the correct story unusually fast generation after a noticeable initial wait—not instant frontier intelligence.
Key facts
- Inception claims 1,107 output tokens per second; Artificial Analysis measures 780.8 through the API.
- Artificial Analysis ranks Mercury 2.5 second for output speed and gives it an Intelligence Index of 12.
- The same service reports a 2.91-second time to first answer token.
- Primary source: Inception's Mercury 2.5 announcement.
Conventional language models produce a token, then the next, from left to right. A diffusion language model starts with an incomplete sequence and repeatedly refines multiple positions in parallel. That is closer to revising a whole line of a draft at once than typing one character after another. Inception says Mercury 2.5 improves intelligence by 40% over Mercury 2 while retaining the serving profile, adds 260K context, tunable reasoning, parallel tool calls and schema-aligned JSON. It remains closed: there is no published checkpoint, download size, parameter count or local VRAM requirement.
The independent numbers come from Artificial Analysis, whose methodology tests a roughly 10,000-token input and at least 1,500 output tokens repeatedly. It measures output speed after the first streamed chunk. For reasoning models that hide some thought tokens, it bases the measure on the last 80% of answer chunks. Time to first answer includes input processing and hidden reasoning before visible output. Those definitions matter: a model can be dazzlingly quick once it starts and still feel slow for a short interactive question.
Mercury's quality evidence is mixed. Artificial Analysis assigns it an Intelligence Index of 12, below the 13 median for similarly priced models, and observes 35 million output tokens versus an 85 million median, suggesting it tends to be concise. Inception's model comparisons are vendor comparisons, not independent proof that the model matches the best generalist systems. The related Flash-dLLM paper supports the general plausibility of diffusion-model systems optimizations but explicitly does not validate Mercury or open-ended dialogue.
Pricing also needs care. Inception announced $0.04/$0.15 per million input/output tokens as a launch price, while AA currently lists $0.25/$0.75 for the observed API endpoint. Treat the active price as an operational check, not a historical marketing claim.
The useful implication is narrow but commercially meaningful. Many agents spend their time on small, structured actions: extracting fields, routing a ticket, generating JSON, rewriting a query or issuing parallel tool calls. In those workloads, throughput can matter as much as leaderboard intelligence. Mercury 2.5 is evidence that diffusion language models have become a shipping inference category. It is not evidence that fast output rate has solved end-to-end latency or quality evaluation.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)