Nvidia Backed the Startup Trying to Make AI Cheaper Than Nvidia Wants It to Be

The AI money has spent two years flooding into training: the eye-watering runs that produce a shiny new model. But the actual bill most companies pay is for inference, the far less glamorous business of running a model millions of times a day once it exists. On 16 July, an infrastructure startup called Fireworks AI raised $1.5 billion in a Series D at a $17.5 billion valuation, betting the whole future is in that second, boring number.

Quick definition, because the jargon does real work here. Inference is what happens every time you send a prompt and a model answers. Training is the one-off capital cost. Inference is the metered, forever cost, the electricity bill of AI. Fireworks runs open-weight models (the ones whose innards are published, so anyone can host them) fast and cheap, and sells that as a service to companies that do not fancy renting a frontier model at frontier prices.

Who wrote the cheques, and why it is spicy

The round was led by Atreides Management, Index Ventures and TCV, with a long tail of names including Lightspeed, Bessemer, Menlo, Insight, Lone Pine and the Ontario Teachers’ Pension Plan. The one that makes you raise an eyebrow is Nvidia. Nvidia’s entire fortune rests on people spending enormous sums on its chips. Fireworks exists to squeeze more output from every chip, so buyers spend less per answer. Nvidia backing it is either a hedge, a land-grab, or a bet that cheaper inference just means everyone runs more of it. Probably all three.

The traction is not vapour. Fireworks says it crossed $1 billion in annualised revenue, up roughly fivefold on the year, and that daily traffic on its platform grew from about 15 trillion tokens to more than 40 trillion. A token is the chunk of text a model reads or writes, roughly a word-part, and it is the unit these firms bill on. Forty trillion a day is a genuine utility-scale operation, not a demo.

The angle: the value is leaking downhill

Here is the pattern worth clocking. The headlines go to the labs building trillion-pound models. The durable business is turning out to be the layer underneath: whoever runs those models cheapest, fastest and most reliably. Open-weight models keep getting good enough that a lot of companies no longer need the most expensive option, and the firm that hosts the cheaper one efficiently gets to keep the margin. That is a less thrilling story than artificial general intelligence, and it is where a chunk of the actual profit is settling.

The risk, and it is a real one, is that inference hosting is a commodity. Everyone from the big clouds to a dozen well-funded rivals wants this exact business. A $17.5 billion valuation on a service that competes partly on price assumes Fireworks stays meaningfully faster or cheaper than the pack. That is a hard thing to keep proving. But the direction of travel is clear enough: the money is moving from the people who build the models to the people who run them.

Not investment advice. Private valuations are marks set by insiders, not prices you can sell at. This is reporting, not a recommendation.

Did you know: the word “token” that AI firms bill you on has no fixed size. Depending on the model, a single English word can be one token or three, which means two providers quoting the same price per token can charge wildly different amounts for the exact same paragraph.

Related on Top Tool Stack: Nvidia’s $250bn OpenAI backstop · DeepSeek’s free V4 Flash

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →

Sources

Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top