Google's reported Frozen v2 chip bakes model architecture into hardware
Google is reportedly developing an internal server chip called Frozen v2 that embeds the Gemini model architecture directly into silicon. According to sources cited by The Information, the chip could be six to ten times more efficient at serving AI responses than Google's current TPUs, with deployment starting in 2028.
The design hardcodes parts of Gemini's model structure into the hardware, unlike TPUs that flexibly run many models. Because it embeds architecture rather than fixed weights, new weights can still be loaded. The idea reportedly traces to Jeff Dean, chief scientist of Google DeepMind, whose earlier version proposed embedding weights themselves; Google scrapped that because it would only work with a single Gemini version. Frozen v2 is framed as a test run for specialized chips, with smaller production volume than the TPU line.
The efficiency angle is what moves abundance. Serving AI responses is energy- and compute-intensive, and that cost is passed to users. A chip that serves the same model far more efficiently could ease Google's internal compute crunch and let it run models at lower prices, which pushes down the cost of access to capable AI.
The caveats are substantial. The efficiency numbers are projections, not verified results; how much of the architecture will actually be hardcoded has not been decided; and the chip is unlikely to be sold to outside customers. Watch whether 2028 deployment holds and whether real-world efficiency approaches the claimed range.
Source: The Decoder
MANY MINDED