DEV Community

Breach Protocol
Breach Protocol

Posted on • Originally published at groundtruth.day

xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens

xAI released Grok 4.6 on August 12, 2026, and made it available first inside coding tools rather than its own chat product. The company says the model scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol Max and landing one point behind Fable 5 Max at 62. Pricing starts at two dollars per million input tokens and six dollars per million output tokens.

Key facts

  • Announced August 12, 2026 on xAI's news page, positioned as a successor to Grok 4.5.
  • Claimed 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks, versus 56 for Grok 4.5 High.
  • Two dollars per million input tokens, six per million output, with a fast variant at double the price.
  • Available day one in Cursor, Grok Build, the xAI API, and via OpenRouter, Vercel, and Cloudflare.

What xAI says changed

The framing is entirely about duration. "Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work," xAI writes, describing a model that "stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact."

That sentence is a fair summary of where the whole industry has moved. Nobody is selling a better chatbot anymore. The competitive question is how many steps a model can take before it loses the plot, which is why xAI's own benchmark tabs are dominated by agentic coding suites rather than knowledge tests. The company credits longer supplemental training, curated model-generated reasoning data, a revised optimizer and training recipe, and stronger self-testing on long trajectories.

The distribution move is the story

The more interesting decision is where Grok 4.6 shipped first. xAI put it in Cursor and Grok Build on launch day and offered two weeks' worth of usage in one, which is a straightforward bid for the developer seat rather than the consumer one. That is where the recurring revenue is, and it is where a model's step-following ability is measured in public every day by people who will notice immediately if it drifts.

Priced at two dollars in and six dollars out, Grok 4.6 undercuts the top of the market while claiming to sit at it. If the index number holds up in independent testing, that is a real value proposition; if it does not, the price is the concession.

Read the chart carefully

The comparison chart is xAI's own, and xAI notes that "competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards." That is more honest than most vendor charts, but it still means every number on it was selected by the company selling the model, and self-reported competitor figures are not the same as a head-to-head run under identical conditions. The Artificial Analysis Intelligence Index is a composite of nine benchmarks, so a one-point gap is well inside the noise of what benchmark composition choices can produce. Our note on how AI gets benchmarked covers why composite scores compress a lot of disagreement into a single digit.

There is also nothing to download. Grok 4.6 is hosted only, which puts it on the opposite side of a line that DeepSeek crossed this week with an open-weight release of comparable ambition. For anyone whose constraint is data residency, air-gapping, or simply not wanting a vendor to be able to deprecate the model out from under them, a hosted 61 and a downloadable 58 are not competing on the same axis at all.

Why the agentic framing matters

An index score measures a model's answers. What xAI is claiming to have improved is something the index barely captures: whether the model can hold a goal across dozens of tool calls without forgetting what it was doing, second-guessing a decision it already made, or quietly substituting an easier task for the one it was given. That failure mode is the dominant one in production agent work right now, and it is not what a nine-benchmark composite is designed to detect.

The practical test is the one Cursor users will run this week without meaning to. A model that scores well and loses the thread on step forty is worse than a model that scores a point lower and does not. As we noted when the agentic index last shifted, the top of that leaderboard has been separated by single points for months, and single points have stopped predicting which model people actually keep using. xAI's own emphasis on "self-testing on long trajectories" suggests the company knows this too. The claim to watch is not 61. It is whether long-running Grok agents finish.


Originally published on Ground Truth, where every claim is checked against the primary source.

Top comments (0)