Editorial illustration for Google Bakes Part of Gemini AI Directly Into "Frozen v2" Chip
Google Embeds Gemini AI Into Frozen v2 Chip
Google has a chip in the works that doesn't just run Gemini, it partially becomes Gemini. Called "Frozen v2," the internal project reportedly bakes portions of the AI model's architecture directly into silicon, according to sources cited by The Information. Google's current TPU chips are built to serve a wide range of models. Frozen v2 flips that logic, trading flexibility for speed by hardwiring structural elements of Gemini into the hardware itself.
The payoff, according to those same sources, could be efficiency gains of 6 to 10 times over existing TPUs when it comes to serving AI responses. Google isn't planning to roll this out at scale. Deployment is set for 2028, and the company reportedly views Frozen v2 as a smaller-volume experiment rather than a replacement for its main TPU line.
The project traces back to an earlier, more rigid idea from Jeff Dean, Google DeepMind's chief scientist, who first proposed locking a model's actual weights into the chip. That version got scrapped for being too inflexible. What replaced it is a different bet entirely, one that still lets Google swap in new weights while keeping the underlying architecture fixed in hardware.
The chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips, according to sources cited by The Information.
Why this matters
Baking Gemini's architecture into silicon is a bet that model design has stabilized enough to freeze into hardware years before deployment. That's a real gamble: Google is targeting 2028 for Frozen v2, which means locking in assumptions about Gemini's structure now and hoping they still hold when the chip ships. For developers and founders building on Google Cloud, a 6-to-10x efficiency jump would eventually mean cheaper inference and faster responses, but only if Gemini's architecture doesn't change enough to make the frozen portions obsolete before launch.
For researchers, this is a signal that the industry sees architectural churn slowing down, at least for parts of a model worth hard-coding. It also raises a practical question worth watching: what happens to Frozen v2 if Gemini's next major revision changes the very parameters Google chose to freeze. Treat this as an early test of specialized AI silicon, not a preview of what ships next year.
Common Questions Answered
What is the key difference between Google's Frozen v2 chip and current TPU chips?
Frozen v2 bakes portions of Gemini's AI model architecture directly into silicon hardware, whereas current TPU chips are built to serve a wide range of models flexibly. By hardwiring Gemini's structural elements into the chip itself, Frozen v2 trades flexibility for significant speed improvements and efficiency gains.
How much more efficient is the Frozen v2 chip compared to Google's current TPU chips?
According to sources cited by The Information, the Frozen v2 chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips. This substantial efficiency improvement would eventually translate to cheaper inference costs and faster response times for developers and founders using Google Cloud.
What is the main risk Google is taking with the Frozen v2 chip design?
Google is betting that Gemini's model architecture has stabilized enough to permanently freeze into hardware, with the chip targeting 2028 deployment. This is a significant gamble because locking in architectural assumptions now means those design choices must still be optimal years later when the chip actually ships, leaving little room for evolution in the model's structure.
Why would Google choose to hardwire Gemini's architecture into silicon instead of maintaining flexibility?
By hardwiring Gemini's structural elements directly into the Frozen v2 chip's silicon, Google can achieve dramatically improved efficiency and speed compared to general-purpose TPU chips. This specialized approach prioritizes performance optimization for Gemini specifically, accepting reduced flexibility as a trade-off for the 6-to-10x efficiency gains.
Further Reading
- Alphabet stock pops on report it's developing a more advanced AI chip - CNBC
- Google plans new chip to run Gemini models more efficiently - Reuters
- Google develops Frozen v2 server chip to embed Gemini AI - NewsBytes
- Google Reportedly Developing 'Frozen v2' Chip That Hardcodes Gemini Architecture - zglg.work
- Gemini Architecture Written Directly Into Silicon: Technical Details of Google's New AI Inference Chip 'Frozen v2' Revealed - TradingKey