Google’s Gemini 3.7 Flash Lands Cheap and Fast as the Flagship Slips

Three weeks after the last one, Google put out another Gemini. On Wednesday the company released Gemini 3.7 Flash, a fast and cheap “workhorse” model aimed at coding, web development and AI agents, and it undercut its own predecessor’s launch price by half. The cadence is close to the story as the model itself: 3.6 Flash is barely out of the box and 3.7 has replaced it.

The flagship, meanwhile, is still missing. Google has been promising Gemini 3.5 Pro since June and it has yet to appear, so the Flash line sprints out a new version every few weeks while the halo product the company keeps trailing slips further behind, plugged over with smaller, faster models that at least arrive on time.

What is new

Google did not train 3.7 Flash from scratch. The team used algorithmic improvements and user feedback to overhaul the previous model, and the benchmark jumps are chunky for a point release. On the DeepSWE v1.1 coding test the score climbed from 49.0% to 65.3%. On FrontierCode 1.1 Main it went from 34.4% to 43.6%. Those are agent and coding-heavy evaluations, which fits where Google is aiming this thing.

The pricing is the eyebrow-raiser. Gemini 3.7 Flash launches at 75 cents per million input tokens and $3.75 per million output tokens, half the introductory price of 3.6 Flash, though that doubles to $1.50 and $7.50 from January. Cutting the cost of the model while raising the benchmark scores is the sort of move that makes developers pay attention, because inference bills are where AI products live or die.

Built for agents

The framing is deliberate. “Flash” is Google’s label for its efficient models, the ones meant to run at scale rather than win every benchmark crown, and 3.7 is tuned for agent-first workflows where a model makes many calls in a loop rather than answering one clever question. It is live through the Gemini API in Google AI Studio, in Android Studio, in Google’s Antigravity agent platform, and across the Gemini Enterprise Agent Platform and the consumer app.

That spread matters. A cheaper, faster coding model wired into the tools developers already use is how Google turns raw capability into actual usage, and usage is the metric it needs against OpenAI and Anthropic in the developer market.

The flagship shadow

Still, the pattern is hard to ignore. Google is very good at pushing out efficient Flash models on a tight schedule and less good, lately, at landing the flagship that is supposed to headline the generation. Gemini 3.5 Pro was meant to be the halo product. Instead the halo is late and the workhorses are doing the talking.

For most developers that may not matter much. If 3.7 Flash codes better than 3.6 Flash at half the price, the flagship delay is someone else’s problem. But for Google’s positioning against rivals who lead with their biggest models, a top tier that keeps slipping while the cheap tier races ahead is a slightly strange look. Capable, productive, and still waiting on the main event.

Did you know: “tokens” are the chunks of text a model reads and writes, roughly three-quarters of a word each, which is why AI pricing is quoted per million of them rather than per page.

Sources

Related on Top Tool Stack: Grok Bot Gives Every AI Agent Its Own Computer. Here’s the Catch · Oracle Is Putting a 98-Qubit Quantum Computer Inside Its Own AI Data Centre

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →
Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top