Grok 4.7: Twice the Speed, Half the Price — and a Launch Page That Bolds Its Rivals' Wins
SpaceXAI released Grok 4.7 on September 21. The most telling detail is typographic: in the official four-column comparison table, two of the bolded best-in-row scores belong to competitors. CursorBench 4.0 bolds Fable 5.1 Max at 51.8% (Grok 4.7: 46.3%). DeepSWE v1.1 bolds GPT-5.6 Sol at 72.7% (Grok: 71.0%).
A lab printing its rivals' wins on its own launch page is not modesty. It is positioning. The official line reads "highly competitive in its class" — and the class it chose is price.
Built for multi-hour tasks
Per the launch page: a new, larger base model than Grok 4.6, a longer reinforcement-learning run on a harder task mix "weighted toward problems that take many hours to complete," better self-verification and long-context management. It was also trained to natively understand the Grok Bot harness — the first time a lab has put a specific harness into the training objective.
Day-one availability: Cursor, Grok Build, the API at console.x.ai, plus third-party harnesses, routers, and clouds. Price and serving speed are unchanged from 4.6: $2 per million input tokens, $6 per million output.
The numbers in context
- CursorBench 4.0: 46.3% (4.6: 40.4%, Sol: 41.7%, Fable 5.1: 51.8%)
- DeepSWE v1.1: 71.0% (Sol: 72.7%, Fable: 70.0%)
- EEBench (electrical engineering): 64.0% — a clear win, 8 points ahead of second place
- Terminal-Bench 4.0: 38.0% — exactly between the frontier tier (GPT-6 Astra 57.9, Opus 5 51.8) and the Chinese flash tier (GLM-5.3-Flash 33, DeepSeek-V4.1-Flash 27)
- Harvey legal agent benchmark: 19.6% — a weak score, printed anyway
Honest scorecard, unusual for a launch page.
Not the crown. The workhorse slot.
Against Fable 5.1 Max, Grok 4.7 delivers roughly 90% of the CursorBench score at one-fifth to one-eighth of the price ($2/$6 vs $10/$50). The most-shared community verdict: third place in agentic coding behind Anthropic and OpenAI, but factoring speed and cost, a great everyday workhorse.
And one genuinely new pricing move: a Fast variant at twice the price for twice the output speed. Speed itself is now a SKU. Cached-input pricing went from one lab's experiment to an industry standard; expect the same here.
Two reversals worth logging
- The famously permissive lab now claims its strongest refusal and jailbreak resistance to date, with a new safeguard stack: HackerBench v0.3 lets only 3.3% of risky dual-use prompts through. Even the rebel lab now sells "containment" as a feature.
- Pre-launch talk vs. delivery: the 2.1T-parameter figure, the "decades of SpaceX data," and the September 12 date all came from pre-launch coverage. None of it appears on the official page, and the model landed nine days late.
My read
I run an LLM gateway (Flatkey), so model pricing is literally my business — read the following with that disclosure in mind.
The third-place strategy is a gift to buyers. The frontier tier finally has an explicitly priced value lane, and not every task deserves crown pricing. On our own gateway, a scenario-based routing regression on real traffic cut model spend by 48.2%. A 90%-capability model at 20% of the price is exactly the slot a router wants to fill.
All benchmark figures are from SpaceXAI's launch page as of September 21, 2026. The community cost comparison ($2.7K for Grok Build vs $3.8K for Claude Code + Fable 5.1 on Artificial Analysis' coding agent index) circulated on X and was not independently verified by me.
Top comments (0)