The models are arriving faster than anyone can use them, and every single launch shows up with the subtlety of a car alarm. This week alone we got two new models from Anthropic, a new one from Google DeepMind, and from OpenAI an announcement about an announcement. And I’ll say the thing you’re not supposed to say: the pace hasn’t slowed, but the part where any of this changes your actual life has been on a break for months.
For most people, ChatGPT and Claude have been good enough for a long time. The real leaps now land in maths, science and coding, which means the people genuinely feeling the difference are coders, and even they’re mostly getting there with a little less prompting rather than witnessing a miracle. Meanwhile the internet insists every release changes everything. So let’s look at what actually shipped this week, minus the confetti.
Fable 5.1: the smartest model on Earth, and it will empty your account
Anthropic’s Fable 5.1 (its stablemate Mythos 5.1 went only to a small set of cyber-defence firms) is what you get on a paid Anthropic plan now. On the benchmarks Anthropic chose to show, it’s a monster: roughly double any predecessor on agentic scientific research, a real jump over GPT-5.6 Sol on agentic coding, a solid bump on knowledge work. Of course they cherry-picked the charts that made 5.6 Sol look anaemic. On the neutral aggregate, Artificial Analysis, it posts a 66, ahead of the previous best, Claude Opus 5 at 63. A three-point gap is genuinely meaningful; that’s about the distance from GLM 5.3 to Opus 5.
It also refuses less. Anthropic tightened its cyber-security safeguards so the model now helps identify software vulnerabilities, with roughly 60% fewer interventions, though it still punts certain dual-use tasks to the Opus models. And they’ve bolted on anti-distillation measures: new API accounts can no longer edit Claude’s prior context in a multi-turn chat while keeping its earlier thinking, which is designed to stop rivals from extracting Claude’s reasoning to train their own models. There’s a rich irony here, given every frontier lab was built by scraping the entire internet, but that’s a rant for another day.
Now the catch, and it’s a big one. Fable 5.1 is priced identically to Fable 5, at $10 per million input tokens and $50 per million output. Anthropic claims 5.1 runs about 25% cheaper for typical workloads thanks to cheaper cache reads. The independent numbers say otherwise: on a cost-per-task basis it’s the most expensive model on the board, roughly $3.69 per task versus $3.14 for Fable 5. It is the single most intelligent model available, and also the single most expensive.
I put it to work. It generated the best SVG of Gary Busey the leaderboard has ever ranked, first place, undisputed. That one picture cost $4.35 and eighteen minutes. It also one-shotted the best clone of the game Mega Bonk I’ve seen, detailed characters, a shop, working menus, the lot. Then I checked the meter: $114 and climbing, because I’d let Claude Code grind on it for over an hour and a half on the $200 “20x” plan, blow through the whole allowance, and start eating overage credits. Superb work. Absolutely ruinous.
Gemini 3.8 Flash: the one nobody’s hyping, which is exactly why you should care
While everyone screamed about Fable, Google released Gemini 3.8 Flash to a shrug. Look at the actual coding benchmark, Deep SWE, and it scores about 74%, tying Claude Opus 5 and edging out Fable 5, the model that was considered too powerful to fully release six weeks ago. The kicker is the price: around $2.36 per task against Opus 5’s $11.84 and Fable 5’s eye-watering $2,163. Per intelligence task it’s roughly 58 cents, versus GPT-5.6’s 95 cents and Fable 5.1’s $3.69. It averages 2.5 minutes a task where Opus 5 and Fable take 7.4.
It is not the smartest model in the room. On the general intelligence index it sits mid-table, around 59, weaker on broad knowledge work. But for coding specifically it’s near state-of-the-art at a tenth of the cost, and my one-shot Mega Bonk from it, while not as pretty as Fable’s, was miles ahead of the sad brown potato Fable 5 handed me on its first try. Google spent so long being the punchline that nobody noticed it strolled back in with the best-value coding model on the table.
The numbers, side by side
| Model | Coding (Deep SWE) | Cost per task | Speed |
|---|---|---|---|
| Fable 5.1 | not yet listed | ~$3.69 (dearest) | ~7.4 min |
| Fable 5 | ~70% | ~$3.14 (Deep SWE ~$2,163) | ~7.4 min |
| Gemini 3.8 Flash | ~74% (ties Opus 5) | ~$0.58 | ~2.5 min |
| Claude Opus 5 | ~74% | ~$2.34 (Deep SWE ~$11.84) | ~7.4 min |
| GPT-5.6 | strong | ~$0.95 | ~3.9 min |
OpenAI Astra: the sequel where the AI learns to hack
This one isn’t a release, it’s OpenAI clearing its throat. In a “path to Astra” note, OpenAI says it’s been holding the model back on safety grounds, because with the right tools it can find previously unknown security flaws and build exploits across well-protected systems with nobody guiding each step. It’s the first model they’ve designated at this threat level, and they’ve delayed parts of it to harden the guardrails. “Available soon,” they say, and it’s hard to think of a less soothing sentence than a company announcing it built something too dangerous to ship, back shortly.
The efficiency jump is the eye-opener. On OpenAI’s own chart, their current best, GPT-5.6 Sol, hit an 11.5% success rate at finding and exploiting vulnerabilities using almost 140,000 tokens. Astra reportedly hits 17.5% using just 18,535 tokens, and climbs toward 40% with 76,188. So it got both far more effective and far more efficient at the exact task you’d least want it efficient at.
There’s a deeper worry. Per The Information, Astra uses a training technique called recurrent depth, or a looped transformer, which improves answers by processing the same text repeatedly but obscures the chain of thought. Today you can open a model’s reasoning and audit how it reached an answer. With this, the working happens as unreadable internal maths and you just get the output. For everyday questions, fine. For coding, cyber and biology, losing the ability to see why a brilliant-at-hacking model is heading somewhere is precisely the safety feature you’d want to keep, and it’s the one being switched off. And once the frontier normalises it, open-source and overseas labs with a looser attitude to safety will happily copy the trick.
So what should you actually use?
If you write code for a living, Fable 5.1 is the best there is, and you’ll feel it, right up to the invoice. Reserve it for the genuinely hard problems and keep a hand on the meter. For everyday coding and most real work, Gemini 3.8 Flash or GPT-5.6 gets you to the same place for a fraction of the money; the premium model saves you a couple of prompts, not your afternoon. And if you’re not a coder, none of this changes your Tuesday, whatever the hype machine tells you.
What this means
Two of this week’s three are tools you can use tomorrow, one is a warning shot, and the honest read on the lot is that we’re on a hamster wheel of half-percent gains dressed up as revolutions. Most of that noise comes from creators, not the labs. For your actual workflow: pick the cheapest model that clears your bar, keep one premium model on standby for the hard 10%, and ignore anyone telling you that you need the newest thing by Tuesday. We do not need a new messiah every seven days.
Related on Top Tool Stack: Best AI Coding Assistants in 2026 · ChatGPT vs Claude vs Gemini
Did you know: the same task that costs about 58 cents on Gemini 3.8 Flash can cost north of $11 on a top-tier model for a near-identical result. Most of the AI bill people complain about is a choice, not a law of physics.