2 min read
Here is a problem most companies would happily kill for, and a warning the rest of the industry should heed anyway. Moonshot AI, the Chinese startup behind the Kimi models, got so much demand for its new Kimi K3 that it stopped accepting new subscribers this week, and not because the model flopped, but because it worked: everyone piled in at once and the company ran out of computer to serve them on.
Popularity as a failure mode
Kimi K3 is a 2.8-trillion-parameter open-weight model that earned a reputation fast for chewing through long documents and gnarly tasks cheaply. The reward for that reputation was a stampede, and the stampede hit a wall: not enough servers, accelerators and networking to keep the service reliable. Existing users keep their access; newcomers get a polite “come back later.” In a market where a rival is one browser tab away, “come back later” is precisely how you gift your users to someone else.
The part that is really about chips
This is the US export controls biting, just not in the way the headlines usually frame it. Chinese labs can plainly build models that people want. What they cannot easily buy is a mountain of Nvidia’s best accelerators to serve those models at scale, because Washington has spent two years restricting exactly that. So a Chinese model can win on quality and still lose on availability, throttled not by a shortage of talent but by the size of the compute pipe it is allowed to plug into.
The likely response more or less writes itself: lean harder on domestic Chinese accelerators, squeeze more out of smaller and more efficient models, and ration usage until the servers catch up. It is also a loud advertisement for every company selling inference efficiency, because the lesson of Kimi K3 is that in 2026 a model is only as good as your ability to keep it running when the entire internet shows up at once.
The AI tool stack actually worth paying for
One email a week. The tools, models and moves that matter, minus the hype and the horseshit filter set to maximum. Free.
Did you know: serving a large model to millions of users can cost far more, over its life, than training it did in the first place. Training is a one-off capital splurge; inference is the electricity bill that never stops, which is why “can you afford to run it” is fast becoming a more important question than “can you build it”.