DEV Community

Cover image for Grok 4.7: The Model Is Weak, the Driveway Is the Moat
Max Quimby
Max Quimby

Posted on Originally published at computeleap.com

Grok 4.7: The Model Is Weak, the Driveway Is the Moat

On September 21, xAI shipped Grok 4.7 — a model that burns 125% more output tokens per task than its predecessor while scoring just 2 points higher on the Artificial Analysis Intelligence Index. One day later, Tesla announced that Grok Bot is now live in vehicles, turning every connected Tesla into a voice-driven AI agent that orders your coffee, books your dinner, and manages your email — all hands-free. These two events are not separate stories. They are a single strategy, and if you only read the benchmarks, you will miss it entirely.

📖 Read the full version with charts and embedded sources on ComputeLeap →

The thesis: xAI is not trying to win the model race. It is trying to win the distribution race. And it has a distribution channel that no other AI lab can replicate — millions of cars already parked in driveways.

The Token Efficiency Problem Is Real

Let us start with the numbers, because they are damning.

Artificial Analysis benchmarked Grok 4.7 at its xhigh reasoning effort and found it consumes approximately 81,000 output tokens per Intelligence Index task. Grok 4.6 at its high setting used 36,000 tokens for the same work. That is a 125% increase. GPT-6 Astra at max uses 27,000 — making Grok 4.7 nearly three times as verbose for comparable quality.

Artificial Analysis benchmarking Grok 4.7 showing token consumption and intelligence index scores

View original analysis on Artificial Analysis →

The rate card did not change: $2 per million input tokens, $6 per million output. But the cost per task tells the real story. Grok 4.6 cost $1.86 per Intelligence Index task. Grok 4.7 costs $3.74 — more expensive than GPT-6 Astra at $3.26, despite Astra scoring 7 points higher on the Intelligence Index (53 vs 46).

⚠️ Rate card vs. reality: Grok 4.7 looks cheaper than GPT-6 Astra on the price sheet ($6 vs $10 per million output tokens). But it burns 3x the tokens per task. The real cost is $3.74 per task vs Astra's $3.26. The sticker price is marketing. Token consumption is engineering.

As we covered in our analysis of AI token economics, the sticker price is increasingly disconnected from the real cost of getting work done.

Elon Promised Better. Theo Browne Called It Out.

What makes this worse is that Elon Musk specifically promised Grok 4.7 would be more token-efficient. Developer and YouTuber Theo Browne delivered one of his most confrontational videos in months, pulling the receipts: "I have yet to find a single bench where Grok 4.7 is more token-efficient than Grok 4.6." The regression is not 5% or 10% — it is 30 to 80%, depending on the benchmark.

Theo's analysis exposed another layer of misleading framing. Cursor CEO Michael Truel pushed back claiming only "5% more tokens for median requests." Theo correctly identified the distinction: if the model issues twice as many requests per prompt — which is exactly what deeper chain-of-thought does — then "5% more per request" can be technically true while the total bill doubles. The per-request claim is the kind of stat designed to survive a tweet but collapse under a spreadsheet.

Theo's strategic read is worth noting: xAI likely kept running reinforcement learning on the old base after the new 2.1-trillion-parameter base underperformed, spending a week of RL making the model reason longer and verify more. That is precisely where the token efficiency went — the model got better by talking more, not by thinking more efficiently.

What Grok 4.7 Actually Does Well

Credit where it is due: Grok 4.7 is not uniformly bad.

On Terminal-Bench 4.0 — which measures multi-hour terminal tasks — it nearly doubles its predecessor, jumping from 20.3% to 38.0%. On DeepSWE v1.1, it scores 71.0% versus Claude Fable 5.1 Max's 65.2%. On the Artificial Analysis Coding Agent Index (with Grok Build), it climbs from 47 to 56, ranking fourth behind Fable 5.1, GPT-6 Astra, and Claude Opus 5.

These are real gains in the long-running, multi-step tasks that matter for agentic use cases. But they come at a cost that the rate card hides, and on the overall Intelligence Index — the broadest measure of general capability — Grok 4.7 lands mid-pack at 46, well behind the leaders at 53.

As The Decoder put it: bargain prices, wide gap. Or as VentureBeat framed it: coding gains paired with token consumption that threatens real-world ROI.

Meanwhile, the Driveway Strategy Ships

Here is where the story splits in two — and where most analysis stops too early.

On September 22, one day after the model launch, Tesla began rolling out Grok Bot integration to vehicles. This is not a concept or a roadmap item. It is live. Drivers can review incoming email, manage calendars, place DoorDash orders, book OpenTable reservations, schedule appointments, and delegate multi-step errands — all by voice, while driving.

Elon Musk on X announcing Grok Bot is now available in Tesla vehicles

View original post on X →

Elon was not subtle about the play. "Grok Bot usage is growing faster than anything we've ever seen," he posted, in a tweet that hit 10.3 million views and nearly 12,000 likes. The framing was not about benchmark scores. It was about adoption velocity.

Elon Musk on X: Grok Bot usage is growing faster than anything we have ever seen

View original post on X →

Alex Finn's hands-on video from his Cybertruck captures the user experience side of this bet. He called it "the coolest AI integration ever" — and whether or not you agree with the hyperbole, the demo is real: voice-driven task execution, native vehicle integration, no app switching. This is not a chatbot. It is an embedded AI agent with access to your car, your calendar, and your credit card.

The Distribution Thesis: Why the Model Might Not Matter

Here is the uncomfortable insight that the benchmark fixation misses.

Most AI users will never run a benchmark. They will never compare token costs per task. They will never read an Artificial Analysis chart. They will talk to whatever AI is already in their environment — their phone, their car, their social feed — and they will judge it on whether it gets their coffee order right.

Tesla has an install base that no AI lab can replicate. Over-the-air software updates can deliver a new model to millions of connected vehicles simultaneously. The car is already in the driveway. The user is already locked into the Tesla ecosystem for navigation, charging, and increasingly, for software features they pay a subscription for.

This is the Amazon Echo strategy applied to cars: the hardware is the distribution channel, the AI is the lock-in, and the quality only needs to be "good enough" — not best-in-class. Alexa was never the smartest assistant. It was the one in 100 million kitchens.

💡 xAI's distribution flywheel: X (hundreds of millions of users) + Tesla (millions of vehicles) + Grok Bot (agentic capabilities) + eventually Optimus (physical robots). The model is the commodity. The surface area is the moat.

What the Community Is Saying

The Hacker News discussion tells a story of bifurcated sentiment. Readers zeroed in on the intelligence-to-token chart and called the upgrade a regression in efficiency. A separate thread highlighted that Grok 4.7 carries 40% more weights than 4.6 at the same price — suggesting xAI is eating the inference cost increase rather than passing it through, a move that only works if they expect distribution volume to subsidize quality.

Hacker News discussion on Grok 4.7 with community commentary on token efficiency regression

View discussion on Hacker News →

On X, the reaction split cleanly between the AI research community (skeptical of the benchmarks) and the Tesla community (excited about Grok Bot). Elon's own framing was telling: he called 4.7 "fairly well for a smallish model" — lowering the bar in a way that would be unthinkable from Anthropic or OpenAI.

Elon Musk on X: Interesting. Grok 4.7 is performing fairly well for a smallish model.

View original post on X →

When your CEO describes your flagship release as "fairly well," the product is not the point. The platform is.

Teslaconomics, a prominent Tesla-focused account, captured the enthusiasm gap: "This is truly one of those things that's hard to fully understand the power of" — the language of ecosystem believers, not benchmark readers.

Teslaconomics on X: Grok Bot is officially available in Tesla vehicles with early access integration

View original post on X →

The Contrarian Corner: What if Distribution Wins?

⚠️ Here is the case against the conventional "Grok is falling behind" narrative: in consumer AI, distribution has always beaten quality. Google Search was not the first search engine. Chrome was not the first browser. Android was not the best mobile OS. They won because they were everywhere.

If xAI gets Grok Bot into 10+ million Teslas, 500+ million X accounts, and eventually into Optimus robots — the question is not whether Grok 4.7 beats Fable 5.1 on MATH-500. The question is whether most users will ever encounter Fable 5.1 at all.

The counter-counter: distribution without quality creates churn. Alexa is in 100 million kitchens and users barely talk to it anymore. Distribution gets you trial. Quality gets you retention. xAI needs both, and right now they are only shipping one.

What This Means for Builders

If you are consuming AI via API: Do not trust rate cards. Grok 4.7 at $2/$6 per million tokens looks cheaper than GPT-6 Astra until you realize it burns 3x the tokens. Profile your actual workloads. Cost per task, not cost per token, is the metric that hits your budget. We broke this down in detail in our token economics deep dive.

If you are building AI products: xAI's distribution play is a thesis you need to grapple with. If you are building a standalone AI product that requires users to seek you out, you are competing against AI that is already in their car, their social feed, and eventually their home robot. Either build where the users already are (platform integrations) or build something so differentiated that they will switch.

If you are evaluating model providers: Grok 4.7 has a niche. Long-running agentic coding tasks, multi-hour terminal sessions, and legal workflows (it leads on the Harvey Legal Agent benchmark) are genuine strengths. But for general-purpose intelligence and cost-efficient inference, Claude and GPT-6 remain the leaders. The frontier release war continues to be a three-horse race where xAI is the value pick with caveats.

The Bottom Line

Grok 4.7 is a mediocre model release wrapped around a brilliant distribution play. The model burns more tokens, costs more per task, and scores lower on general intelligence than its competitors. But it shipped the same week that Grok Bot went live in Teslas — and that timing is not coincidence.

xAI is not trying to win the benchmark war. It is trying to make the benchmark war irrelevant. Whether that bet pays off depends on a question no leaderboard can answer: when AI is already in the driveway, does anyone bother shopping for a better one?


Originally published at ComputeLeap

Top comments (0)