OpenAI Cuts the Price of Its Best Model, and the Fine Print Does the Talking

On 22 August, OpenAI trimmed the price of GPT-5.6 Sol, its flagship model. Standard input tokens fell 20% to $4 per million, and output dropped 33.3% to $20 per million. For anyone running a product on top of Sol, an output job that cost $30 last week now costs $20. That is real money once you are feeding the meter millions of tokens a day.

Read the fine print and the generosity gets conditional. OpenAI calls the new $4/$20 rate “promotional” and only guarantees it through 21 November. The pricing trackers cannot even agree it happened: AI Pricing Guru logged the cut on 22 August, while BenchLM’s page the same day still listed Sol at $5/$30. When the people whose whole job is watching AI prices are out of step, you are the one left checking the invoice.

A price war nobody can afford to lose

The flagship tier has bunched up. Sol’s new $4/$20 sits a shade under Anthropic’s Claude Opus 4.8 at $5/$25, OpenAI’s mid-tier Terra runs $2/$12 against Claude Sonnet 5, and Anthropic’s priciest model, Fable 5, lists at $10/$50. With headline numbers this close, a 20% chop is a shove, not a rounding error.

It also lands while the money at stake is enormous. Anthropic’s annual revenue run rate topped $65bn by the end of July, up from about $9bn at the end of 2025, according to Bloomberg. OpenAI is chasing the same enterprise budgets. Cutting list prices into that fight is a monetisation gamble, because the margins underneath are thinner than software investors are used to. ICONIQ’s 2026 snapshot pegs the average AI product gross margin at 52% for this year, up from 41% in 2024, but still miles below the 80%-plus that classic software enjoys.

Cheaper tokens, bigger bills

The twist that catches finance teams out is this: per-token prices are falling, yet total AI bills are rocketing, because usage is climbing faster than prices drop. Ramp’s spending data shows the top 1% of US companies hit a median of $7,400 per employee per month on AI in July, against $11.95 for the median firm, a gap of more than 600 to 1. Spend per employee has more than tripled across the board in a matter of months.

Reasoning models and agent loops are the culprits. A single agentic task can chew through millions of tokens as the model thinks, retries, and calls tools. A 20% cut on the input rate is cold comfort if your agent is now doing ten times the work it did in the spring.

How to actually pay less

The real savings hide in the boring settings. Cached input on Sol bills at 10% of the standard rate, so a reused system prompt costs $0.40 per million instead of $4. The Batch API takes a flat 50% off input and output for jobs that can run asynchronously. And routing routine work to GPT-5.6 Luna, at $0.20/$1.20, rather than defaulting everything to the flagship, is the biggest lever most teams never pull. Stack a cached, batched Luna workload and you are paying a fraction of the sticker price splashed across the front page.

None of that turns up in a press release about a headline price cut. But it is where the money actually is, and it is the part you control.

Did you know: OpenAI’s cheapest current model, GPT-5.6 Luna, costs $0.20 per million input tokens. Its priciest, GPT-5.5 Pro, costs $30, a 150-fold spread inside a single company’s own menu.

Sources

Related on Top Tool Stack: Rillet Hit a $1bn Valuation in Under 48 Hours, and NetSuite Should Be Nervous · Amazon, Google and Microsoft Are Spending 102% of Cloud Revenue on AI Kit

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →
Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top