
In a shock to precisely no one, when you let the tech gods grade their own work the result is less an evaluation than a public round of self-applause. Exhibit A: Google’s new Gemini 4 Argon, which, in Google’s own benchmark table, conveniently beats GPT-6 Astra on 14 of 19 tests, Claude Opus 5.5 included. Hand the marking to someone neutral and the halo slips. The one figure that genuinely stands out is the hallucination rate, where Argon invents things 15 percent of the time against Astra’s 51, which raises the question of what OpenAI has been feeding its model and whether it was prescribed. And the punchline: Argon is barely available, out first to a handful of cyber-defence testers, so you cannot even take it for a spin to check Google’s maths.
What Google says
Gemini 4 Argon landed on 30 September as Google DeepMind’s new frontier model, with output capacity jumping to one million tokens from the previous 64,000. In Google’s own 19-benchmark comparison it tops GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on most tests. On DeepSWE v1.1, a software-engineering benchmark, Argon scores 77.9 percent against 74.2 for Opus 5.5 and 74.1 for Astra. On AutomationBench it posts 51.3 percent versus 42.5 for Opus 5.5 and 41.4 for Astra. On the long-video benchmark LVBench it reaches 91.7 percent, ahead of Astra at 87.5 and Opus 5.5 at 83.7.
What the neutral umpire says
Independent lab Artificial Analysis is less dazzled. On its Intelligence Index, Argon scores 53, level with GPT-6 Astra but clearly behind Anthropic’s Opus 5.5. In other words, Google’s “beats everyone” table and the independent “matches one, trails the other” result are both true, depending on who holds the red pen. For a buyer, that gap between a vendor’s slide and a third party’s test is the whole lesson.
The one number worth keeping
Argon’s standout is reliability, not raw intelligence. Its 15 percent hallucination rate is the lowest of any model in its tier, against 51 percent for GPT-6 Astra and 54 percent for GPT-6.1 Sol. If those figures hold up in daily use, a model that makes things up a third as often as its rivals is a bigger deal for real work than a point or two on a coding leaderboard.
The catch
You cannot have it yet. Argon is rolling out through a limited programme called Fairwind, starting with cyber-defence testers rather than the general public, so independent verification and real-world stress-testing are both still pending. Until then, Google’s benchmark table is a claim, not a result.
What this means
Treat any model’s self-reported benchmark table as marketing until a neutral party repeats it. On current independent numbers Opus 5.5 still leads on intelligence, Argon and Astra are roughly matched, and Argon’s low hallucination rate is the genuinely interesting development worth tracking once anyone outside Google can actually use it.
Sources: Google DeepMind (30 Sept comparison table); VentureBeat; The Decoder; Artificial Analysis Intelligence Index.