3 min read
Google is reportedly building a chip that hardwires its Gemini AI model directly into silicon, and if the numbers hold, it could be six to ten times more efficient than the TPUs Google runs today. The chip, known internally as Frozen v2 and first reported by The Information, is exactly the kind of quiet infrastructure move that decides who can afford to serve AI at scale. Which is to say it matters more than most of the model launches that get bigger headlines.
What “baking the model into the chip” means
Normal AI chips are generalists. They load a model into memory and shuffle enormous amounts of data back and forth to do the maths, which is why they cost the earth and burn power like a small town. Frozen v2’s idea is to hardcode parts of Gemini’s neural-network blueprint straight into the physical circuitry, so the hardware performs fewer calculations and moves data far shorter distances. Less shuffling means less power and more tokens per watt. The claimed 6-to-10-times efficiency on power-per-token would be the largest single-generation jump in Google’s custom-silicon programme.
Why Google, and why now
The timing is telling. Google has missed its Gemini 3.5 Pro target three times and just been ordered by the EU to open Android to rival AI assistants. A chip that cuts serving costs by most of an order of magnitude would let Google compete hard on price even while its flagship model lags, which is precisely the sort of structural advantage that survives a bad model quarter. Custom silicon is the one area where Google’s decade-long head start is undisputed, and it is being pushed by a real internal compute crunch: Google Cloud has reportedly been turning away outside deals for lack of capacity.
The obligatory cold water
Now the skepticism, because it is warranted. This is reported from internal sources, not announced, and efficiency claims made before a chip reaches production are marketing-adjacent by nature. The 6-to-10-times range is wide enough to contain wildly different outcomes, efficiency depends heavily on the workload, and the release is slated for 2028 with Google reportedly treating it as an exploratory project it may never build at the scale of its TPU fleet. Real direction, unproven magnitude, years away.
The pattern worth watching
Frozen v2 is the same bet the startup Etched is making with its transformer-only chips, which we covered recently: give up flexibility, hardcode one architecture, and win enormously on speed and cost if that architecture stays king. When the company that helped invent the modern AI chip starts doing this too, the industry is telling you where it now thinks the efficiency gains live. Not in bigger models, but in silicon shaped around the model you already have.
The AI tool stack actually worth paying for
One email a week. The tools, models and moves that matter, minus the hype and the horseshit filter set to maximum. Free.
Did you know: Google’s TPUs, its custom AI chips, have quietly underpinned its AI ambitions since 2015, long before the current boom, and they are the main reason Google can serve models without buying its entire fleet from Nvidia like everyone else. Frozen v2 would push that home-grown advantage further, which is exactly why a chip nobody can buy until 2028 is worth paying attention to now.