Alibaba’s Qwen3.8-Max: a 2.4-trillion-parameter frontier model, and the weights are coming

Alibaba put out its biggest model yet on Monday 3 August 2026, and the headline number is not the interesting part. The interesting part is the price tag, and the promise to hand you the weights.

Qwen3.8-Max is a 2.4 trillion parameter model built on a mixture-of-experts design. That phrase matters, so here is the plain version: mixture-of-experts (MoE) means only a slice of that giant network switches on for any given question, so you get the reach of an enormous brain without paying enormous-brain compute on every single query. It handles text, images and long-form video, and it takes a context window of up to one million tokens, which is roughly 750,000 words in one go.

Where you can get it, and for how much

It is live globally through Alibaba Cloud’s Model Studio API and through QwenWork, the company’s workplace agent platform. Alibaba priced it at about 40% of Claude Opus 5 for input tokens and around 24% for output in international markets. Then the kicker: Alibaba said it would release the model’s weights for public download the following week, marking a return to the open strategy it had drifted away from on several recent flagship releases.

For anyone building on top of these models, the cost floor just dropped again, and this time it dropped from a lab that says it will let you download and fine-tune the thing yourself.

The bit the press releases skate over

Alibaba claims benchmark scores that rival the best from Anthropic and OpenAI. Claims. Vendor benchmarks are marketing until an independent party reruns them, so treat the leaderboard talk as a starting bid, not a verdict. At the time of writing the open weights were promised, not published, so the download you actually care about was still a week out.

And the obvious caution: this is a Chinese-hosted frontier model. If you are moving regulated or sensitive data, the API endpoint and the jurisdiction matter as much as the token price. Weigh that before you wire your production stack to it.

Still, the direction of travel is the real story. The frontier is no longer three American labs renting you access through a turnstile. A Chinese lab is undercutting them on price by more than half and going open at the top end. If you are a small team that wants to run or customise a serious model without asking a gatekeeper’s permission, that is a genuinely good week, whatever you make of the geopolitics.

Hong Kong-listed Alibaba shares jumped in early trading after the announcement, which tells you the market read this as a shot at the incumbents rather than a science-fair entry.

Did you know: despite carrying 2.4 trillion parameters, an MoE model only fires a fraction of them per request, which is exactly why a model this big can be sold this cheap.

Related on Top Tool Stack: Moonshot’s Kimi K3 open weights · DeepSeek’s free V4 Flash

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →

Sources

Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top