Google Rolls Out Cheap Gemini Flash While the Flagship It Promised in June Is Still Missing

Google put out Gemini 3.6 Flash on 21 July, and on its own terms it is a tidy update. Flash is the cheap, fast tier, the one you reach for when you want good-enough answers at volume rather than the smartest possible reply. The new version uses about 17% fewer output tokens than the previous 3.5 Flash to say the same thing, and Google cut the output price to $7.50 per million tokens from $9.00, while holding input at $1.50. Fewer tokens and a lower rate is a real saving if you run anything at scale.

It arrived day one across AI Studio, the Gemini API and the Gemini app, with the usual million-token context window. Google also put out Gemini 3.5 Flash-Lite, an even cheaper tier, and a locked-down ‘Flash Cyber’ variant tuned for security work and handed only to governments and trusted partners.

The elephant not in the room

Here is the bit Google would rather you skimmed past. There is still no Gemini 3.5 Pro. That is the flagship, the model meant to go toe to toe with GPT-5.6 and Claude Opus 5. Sundar Pichai stood on stage at Google I/O in May and told developers it would land in June. It is late July and it is nowhere. Google now says the Pro model fell short of its own internal targets on coding and complex reasoning, so the wider release slipped, with no new date attached.

So the pattern is a run of cheap Flash models out the door on time, and the one that actually has to beat the competition kept in the workshop. Delaying a model that is not ready is the responsible call. Announcing a June date to a room full of developers building around your roadmap and then going quiet is a good deal less responsible.

Meanwhile, a promise about the next one

On the same day, Google said it had begun what it calls its most ambitious pretraining run yet, for Gemini 4. Read that back. The June flagship has not turned up, and the pitch has already moved to the version after next. For anyone keeping score, promises about future models are cheap, and Google is not short of them.

If you build things for a living, ignore the roadmap theatre and judge Flash on what it is: a cheaper, leaner workhorse that is genuinely good value for high-volume, low-stakes work. Just do not plan your product around a Pro model that keeps missing its own bus.

Did you know: ‘Flash’ tiers exist because most real-world AI tasks (summarising, tagging, simple chat) do not need a genius, and paying flagship rates for them is like hiring a barrister to read your post.

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →

Sources

Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top