
Two of the biggest names in AI put out new models within a week of each other, and both would like you to believe theirs is the sensible financial choice. xAI released Grok 4.7 on 21 September; Anthropic answered with Claude Sonnet 5.5 on the 28th. Line up the sticker prices and Grok wins in a walk: six dollars per million output tokens against Claude’s ten. If that were the whole story, this would be a very short article.
It is not the whole story, because nobody prices AI honestly on the first line anymore.
The catch lives in the cache
Most real work with these models reuses the same big lump of context on every call. A coding agent re-reads your codebase. A support bot re-reads the same policy documents. What matters there is not the headline input rate but the cached-read rate, and this is where the “cheap” model develops an asterisk. Grok 4.7 discounts cached reads by 75 percent. Claude Sonnet 5.5 discounts them by 90, dropping to twenty cents per million against Grok’s fifty. Run a workload that leans on cached context, which is to say most agent workloads, and the model with the higher headline price ends up billing you less.
Grok has two more asterisks. Its output rate climbs to twelve dollars once a single request crosses 200,000 input tokens, so the long jobs cost more precisely when you can least avoid them. And its context window tops out at 500,000 tokens against Claude’s full million. On raw benchmarks the two trade blows, with Sonnet 5.5 leading on the coding and agentic evaluations most businesses actually care about. So the “obvious budget pick” is only obvious if you never read past the first number.
Everyone is racing to the bottom on purpose
Zoom out and the individual price tags matter less than the direction of travel, which is straight down. Anthropic, OpenAI and xAI have spent the last month undercutting each other by the hour, and it is not because running these models got dramatically cheaper overnight. It is because IPO season is coming and market share is the number that impresses bankers. Get the developers hooked, get the enterprises to build their entire stack on your API, book the growth chart, then go public.
You have seen this film before. It was called streaming.
The streaming playbook, now with GPUs
Remember when streaming was cheap, ad-free and good? Netflix ran at a loss for years to hook you, then raised prices more than a dozen times, bolted on an ad tier, and started cancelling shows two seasons in while charging you more to watch whatever survived. Disney+ arrived at seven dollars, swore it would stay affordable, and now sits north of fifteen for a service that pushes ads at you and reruns the same six franchises being strip-mined into oblivion. The quality went down. The price went up. The subscriber, who cancelled cable for exactly this reason, now pays more than cable did and gets less.
The AI price war is that same playbook with a bigger electricity bill. The loss-leader pricing is a customer-acquisition tactic dressed up as generosity. The moment switching costs are high enough and the IPOs have cleared, the enterprise contract that looked like a bargain gets its Netflix moment. So enjoy the twenty-cent cached reads while they last, and design your stack so you can move providers when the inevitable arrives. The one lesson streaming taught everybody is that loyalty is a liability the vendor is counting on.
For now, if your workload is context-heavy and agentic, Sonnet 5.5 is likely the cheaper real-world choice despite the scarier headline number. If you run short, stateless calls and want the lowest raw rate, Grok 4.7 has the edge. Either way, keep a bag packed.