Google is quietly advancing a hardware initiative that could rewrite the rules of the AI chip game. According to a Monday report from tech outlet The Information citing two sources with direct knowledge of the matter, Google is developing a new server chip codenamed “Frozen v2” that directly hardwires the underlying architecture of its Gemini model into silicon, aiming to trade generality for an exponential leap in AI inference efficiency.
The design targets are extremely aggressive: measured by tokens processed per unit of power consumed, Frozen v2’s efficiency is expected to reach 6 to 10 times that of Google’s latest generation of custom AI chips. Sources indicated the chip could be deployed in real-world operations as early as 2028.
Behind this development lies an increasingly severe AI compute shortage within Google. Sources noted that the compute bottleneck has triggered internal tensions over resource allocation, even forcing Google Cloud to turn away collaboration requests from some external customers. Frozen v2 emerged specifically to address this thorny problem—by pre-embedding portions of the Gemini model’s decision-making logic into the chip, dramatically reducing the number of computational steps and data transfers required at runtime, thereby boosting response speed while significantly lowering per-unit compute power consumption.
A Google spokesperson responded in a statement that the company’s teams “continuously research and explore new innovations to deliver the highest performance and efficiency for users and customers,” adding that “not every project reaches production, but this rigorous exploration is core to our full-stack approach.”
The Critical Trade-off: From ‘Full Hardwiring’ to ‘Flexible Hardwiring’
Frozen v2’s core design philosophy stands in stark contrast to the general-purpose AI chips dominating today’s market. Whether Google’s own Tensor Processing Units (TPUs) or Nvidia’s GPUs, both are designed to be compatible with a wide range of AI models, meaning the chip must perform extensive dynamic, real-time decision-making when running a specific model. Frozen v2 takes the opposite approach, permanently etching the Gemini model’s underlying architecture into the chip’s silicon substrate, pre-hardwiring certain decisions in exchange for extreme operational efficiency.
This design concept did not emerge from thin air. The report revealed that Frozen v2’s predecessor was an “original Frozen” proposal led by Jeff Dean, Chief Scientist at Google DeepMind. That approach was even more radical, planning to burn the model weights themselves directly into the chip. However, this would have meant the chip could only serve one specific version of the Gemini model, giving it an unacceptably short lifecycle, and the proposal was ultimately shelved.
Frozen v2 made a critical adjustment on this foundation, pivoting to a “flexible hardwiring” approach—what gets hardwired is the model architecture, not the specific weights. Google indicated the chip can support updates to model weights, the parameter settings that determine how the model responds to queries. This means Frozen v2 chips can be compatible with subsequent versions of Gemini that use the same underlying architecture, preserving the efficiency advantage while significantly extending the chip’s usable lifespan.
Still, this represents a major gamble. Sources noted that the design effectively means Google is betting it will stick with the current Gemini model architecture for the long haul. Google is carefully evaluating exactly how much information to lock into the chip to strike the optimal balance between flexibility and efficiency.
Industry Race: Specialized Inference Chips Take Center Stage
Google’s move reflects a broader trend in AI chip competition shifting from “general-purpose computing” toward “specialized inference.”
As an industry reference point, Canadian chip startup Taalas has adopted a similar design philosophy of hard-coding specific AI models into chips and has raised over $200 million from investors including Quiet Capital and Fidelity. Meanwhile, SambaNova, d-Matrix, and giants like OpenAI and Microsoft are also developing inference chips, though their technical paths differ from Google’s.
Nvidia is likewise doubling down on the inference market. Last December, the company spent $20 billion on a technology licensing deal with inference chip startup Groq to strengthen its competitiveness in this space.
Google explicitly positions Frozen v2 as a new branch outside the TPU product line, not a replacement for its existing custom chip ecosystem. TPUs, like Nvidia GPUs, are general-purpose AI chips compatible with multiple models; Frozen v2 is purpose-built for Gemini, and the two product lines will develop in parallel.
Notably, Google currently has no plans to scale Frozen v2 production to volumes comparable with TPUs. Sources said the relatively limited deployment scale allows Google to bring in external design and manufacturing partners later in the project cycle, without the need for the large-scale, upfront coordination required for major product launches. Additionally, Google views this generation of the product partly as an engineering experiment—allowing engineers to accumulate experience in developing more highly specialized chips, particularly against the backdrop of AI model architectures gradually stabilizing.
Cost Advantages and Commercialization Prospects
Google’s existing custom chip strategy is already showing early results. This year, Google released its eighth-generation TPU and began selling it directly to cloud computing customers, challenging Nvidia’s dominance. Market reports indicate Google has signed a multi-billion-dollar TPU leasing agreement with Meta and is actively pursuing other cloud service provider customers.
Sources said that developing its own chips has already enabled Google to run Gemini models at lower cost. If Frozen v2 successfully reaches mass production, this cost advantage could be further amplified—a 6- to 10-fold efficiency improvement means that processing the same volume of inference requests would require significantly fewer chips, less energy, and less physical space.
For Google, which is building out AI infrastructure on a massive scale, this is not just about technological leadership; it directly impacts the profitability and market competitiveness of its cloud computing business. Against a backdrop of continuously exploding compute demand and supply-side bottlenecks, the specialization path represented by Frozen v2 could become a critical key for tech giants seeking to break through the “compute siege.”