Google is developing Frozen v2, a chip freezing Gemini architecture into silicon for 6–10× tokens/W, deployment as early as 2028, driven by compute shortage.
Google is developing 'Frozen v2,' a custom chip that freezes parts of Gemini’s architecture into silicon, targeting 6–10× more tokens per watt than its newest TPUs. Deployment is planned as early as 2028, driven by a severe AI compute shortage that has forced Google Cloud to turn down outside customers, per @kimmonismus.
Key facts
- Frozen v2 targets 6–10× more tokens per watt.
- Deployment planned as early as 2028.
- Google Cloud turned down customers due to compute shortage.
- Chip freezes Gemini architecture into silicon.
- TPUs reduced Nvidia dependency; Frozen v2 goes further.
Google is preparing a radical shift in AI hardware strategy with a chip internally called 'Frozen v2,' according to a report from @kimmonismus. The chip aims to deliver 6–10× more tokens per watt than Google’s latest TPUs by freezing portions of Gemini’s architecture directly into silicon — effectively making the model's structure immutable at the hardware level.
The motivation is immediate and stark: Google’s AI compute shortage has become severe enough that its Cloud division has turned down outside customers According to @kimmonismus. This scarcity, combined with the soaring cost of running large-scale inference, has pushed Google to explore extreme efficiency measures.
The efficiency–rigidity tradeoff
Frozen v2’s gains come with a significant constraint. Future Gemini models will be able to use the chip only if they retain the same underlying architecture. Any architectural change — a new attention mechanism, different layer counts, or a revised tokenizer — would render the chip obsolete until a new version is fabricated. This is a stark departure from TPUs, which are programmable for different model architectures.
TPUs reduced Google’s dependence on Nvidia. Frozen v2 would go further, tying Gemini’s architecture directly to the silicon running it. The approach mirrors what Apple does with its Neural Engine, where certain ML operations are hardwired, but on a much larger scale — freezing an entire large language model’s core structure.
Timeline and competitive context
Deployment is planned for as early as 2028, a multi-year horizon that suggests Google is willing to accept architectural lock-in for massive efficiency gains. By comparison, Nvidia’s next-generation Blackwell Ultra and Rubin architectures are expected in 2025–2026, and Google’s own TPU v6 (Trillium) was announced in 2024. Frozen v2 would leapfrog both in raw token-per-watt efficiency, but at the cost of flexibility.
The move also signals that Google sees inference cost — not just training cost — as the binding constraint for scaling AI. If Frozen v2 hits its targets, it could dramatically lower the marginal cost of serving Gemini, potentially enabling new use cases that are currently uneconomical. But it also means Google is betting that Gemini’s architecture will remain stable for years — a risky wager in a field where architectural breakthroughs happen quarterly.
What to watch
Watch for Google’s next TPU roadmap update at its I/O conference in May 2026, where it may officially acknowledge Frozen v2 or reveal a related project. Also track whether Gemini’s architecture stabilizes — any major architectural change in Gemini 3 or 4 before 2028 would undercut Frozen v2’s viability.
Originally published on gentic.news
Top comments (0)