The AI world lit up again last night. Alibaba open-sourced Qwen3.8-2.4T-A95B, DeepSeek shipped the official V4 Pro release, and Elon Musk rolled out his newest flagship model: Grok 4.6.
This time, Musk barely gave Grok 4.5 room to breathe.
Grok 4.5 launched on July 16. Twenty-seven days later, Grok 4.6 is already here.
Its composite Intelligence Index climbed from 56 to 61 — tying GPT-5.6 Sol and trailing Fable 5 by a single point. The price didn't move a cent.
Musk summed up the pitch in one line: "Smart, fast, and absurdly good value."
It has the shape of a familiar story: benchmarks tied, price undercut, rivals stung.
61 Points: Into the Top Tier
How good is it, really?
Across Artificial Analysis's nine-benchmark composite, Grok 4.6 High scored 61 on the Intelligence Index, tying GPT-5.6 Sol Max. The leader, Fable 5 Max, sits at 62 — a one-point gap.
Break down the scorecard and Grok 4.6 actually pulls ahead in places. It beat Fable 5 on three benchmarks: GDPVal-AA v2, AA-Briefcase, and Harvey LAB. On Harvey LAB, the legal-use benchmark, it crushed Fable 5 22% to 14.2% — in a category Anthropic has treated as core high-value professional-work territory.
The weak spots are just as clear. On DeepSWE 1.1, Grok 4.6 scored 65.9%, behind GPT-5.6 Sol's 73% and Fable 5's 70%. On Terminal-Bench 3.0, its 26% trails both as well. Complex software engineering and terminal-based tasks remain a gap.
Even so, the overall scorecard is striking. Remember, OpenAI has called GPT-5.6 Sol "the best model we've ever built." Grok 4.6 walked in and matched it on day one.
A score of 61 puts Grok 4.6 squarely in the top tier.
Built for Long-Running Agent Work
This time, Grok 4.6 didn't pour all its effort into acing single questions. It went after something harder: long-running agentic tasks — cross-codebase analysis, self-verification, complex interaction, and visual work.
On SWE-Marathon, a benchmark built for extended tasks, it hit a 31.9% completion rate, 2.5 points above Grok 4.5.
To build that capability, xAI extended training time and had the model practice repeatedly in environments for writing code, optimizing systems, building web pages, and drafting CAD drawings. Grok 4.6 also trained on anonymized Cursor workflow data, learning how developers actually read a project, edit files, and fix errors.
Running longer is one thing. Whether the resulting web page is any good is another.
On Code Arena: WebDev, Grok 4.6 High currently scores 1618, good for 7th place — a solid jump from Grok 4.5's 13th-place finish.
Progress like that is a sign of where the whole field is headed: frontier models are getting genuinely useful at building software, not just discussing it.
For anyone who'd rather skip the raw model tinkering and go straight from an idea to a working product, AI app builders like HappySeeds take a plain-language description and turn it into a shippable app — backend, payments, and all — in minutes.
More Power, Same Price
The most concrete win here is the price. The standard API still runs $2 per million input tokens and $6 per million output tokens — identical to Grok 4.5.
No price increase on the upgrade, plus double usage quota for the first week. Same money, twice the model.
Artificial Analysis ran the numbers a different way, too. Across its nine tests, Grok 4.6 costs an average of $0.84 per completed task; GPT-5.6 Sol costs $1.23. The two models tie on overall score, but Grok 4.6 gets there for roughly a third less.
Fable 5 scores one point higher and costs $3.14 per task on average.
Arthur MacWaters, co-founder and president of Legion Health, looked at Artificial Analysis's comparison chart and the pricing and had nothing but praise.
This isn't xAI's first run at the value play, either. When Grok 4.5 launched, Musk positioned it as Opus-class performance at a fraction of the cost: 1.5 trillion parameters, priced at just $2/$6. That pricing looked absurd at the time.
Now Grok 4.6 pushes the same strategy further — matching the top tier on benchmarks while holding the price exactly where it was.
4.6 Just Shipped. 4.7 Is Already Coming.
Grok 4.6 barely landed, and Musk is already talking about the next one.
Shortly after launch, he teased on X that Grok 4.7 will clearly outperform 4.6, with release expected in three to four weeks. Initial training is already done. The team is now folding a large batch of SpaceX's internal company data into supplementary training — an unusual move, to say the least.
4.6 caught up to the competition. 4.7 is already on the way. And the price will almost certainly stay exactly where it is.
This race is only getting faster.







Top comments (0)