I spent the morning staring at my own token bill, which is not a sentence I ever expected to type, but here we are. There's a name for this now — "tokenmaxxing" — and it's quietly become the most awkward metric in corporate America.
The consumption race got real, fast
The numbers floating around are genuinely weird. Jensen Huang — who has a very direct financial interest in all of this — said he'd be "deeply alarmed" if a $500,000 engineer didn't burn through at least $250,000 worth of tokens. Databricks' Ali Ghodsi singled out one engineer who spent over $7,000 on tokens in a month. Sendbird built a leaderboard tracking everyone's token burn, which is either brilliant gamification or a spectacularly passive-aggressive HR move, I genuinely can't decide. Harvey, the legal AI, went from whatever it was spending to 12 trillion tokens a month — call it a 12x jump in a year.
But the part that tells you where this is heading: Uber blew through its entire annual AI budget in four months. Four. Meta's Adam Mosseri is already talking about per-employee token limits, and Microsoft quietly cancelled Claude Code licenses to herd everyone back under Copilot. The AI spend leaderboard that made headlines in January? Shut down. The bill arrived, and suddenly everyone's rediscovering the concept of "maybe don't burn all the tokens."
Honestly, it feels like watching a teenager get their first credit card. The enthusiasm is real, the spending is real, and the hangover is arriving right on schedule. To be fair to the optimists, part of this is just adoption — 2026 is the year companies actually wired LLMs into daily workflows instead of demoing them. But a lot of it is people not picking the right model. The spread is brutal: top-tier frontier models run 5 to 10x the cost per million tokens of their leaner siblings. Opus 5 sits around 5x what Haiku costs; some new flagships push 10x. If your team's whole AI budget is flowing through one premium model for what is essentially autocomplete, that's not adoption — that's a procurement problem wearing a fancy coat.
The other path: India builds its own rails
Half a world away, India is quietly going the opposite direction. Sarvam AI's co-founder Vivek Raghavan just landed on TIME's 100 most influential AI leaders list — a solid acknowledgment for a company trying to build India's sovereign LLM stack from scratch. The headline number is a 120-billion-parameter open-weight model picked up by the IndiaAI Mission for governance work, powering programs like Citizen Connect and AI4Pragati. Nvidia and HCLTech are backing them, the valuation sits around $1.5 billion, and they've already shipped two models trained on Indian languages — which matters in a country with, you know, hundreds of them.
Right on their heels, voice AI startup Gnani dropped an entire "sovereign AI stack" called Artha: an open-weight foundational model (Evon v3.3) plus an agentic platform (Plexus), aimed at Indian enterprises and public institutions. Two Indian companies, both betting that open weights plus local control beats renting intelligence from California.
From my perspective, the sovereign AI rush is a little overhyped in the slide decks. Nobody really talks about what it costs to host and maintain a 120B model on Indian soil, and "sovereignty" doesn't magically fix compute or talent gaps. But the direction is sound, and it's a healthy counterweight to a market where one or two labs set the default for everyone else.
Tencent's Hy4 is honest about its own flaws
Over in Shenzhen, Tencent previewed Hy4 — an open-weight mixture-of-experts model with 770 billion parameters in total, though only about 49 billion are active for any given request. It's aimed at software engineering, research, and financial analysis, and Tencent plans to bolt it into CodeBuddy and WorkBuddy.
What I actually respect here: Tencent straight-up admitted in the release that the model can take longer than necessary on complex questions and has a tendency to over-verify its own answers. That kind of candor is rare — most model launches read like a press release from the Department of Always-Winning. That said, 770B total is a hefty thing to host even with MoE sparsity, and the "preview" tag isn't decoration. I'd want third-party benchmarks before getting excited, because "beats rivals in internal testing" is a sentence that's lost all meaning this year.
The uncomfortable background: a fake disease name
And then there's this, sitting underneath all the racing — tokens, models, sovereign stacks. Researchers at Lille University Hospital ran a study on "hallucination by proxy," where an LLM proposed differential diagnoses that included a made-up disease name. 44% of the medical residents involved trusted the fictional "neurocadmiumatosis" — yes, that's not a real condition — enough to factor it into their thinking. The residents who caught it had one thing in common: they'd stopped treating the model's output like a peer-reviewed paper.
That's the part that stays with me. We're building leaderboards for token spend while a sizable chunk of trained doctors will still nod along to a hallucinated diagnosis. The frontier models will sort themselves out; the habit of treating AI output as truth is the thing nobody's fixed yet.
Watch your own token burn, don't pay flagship prices for autocomplete, and double-check anything an LLM tells you about your health. Speaking of which, I'm off to figure out why last month's bill looks like it funded a small rocket launch. See you tomorrow.
P.S. — if you've been tracking your own AI spend and feeling slightly nauseous about it, 7x24planning has some genuinely useful planning tools worth a look when you're not mid-panic.

Top comments (0)