Prompt Caching: The Boring Trick That Cuts Your LLM Bill by 90%
AI on a Budget. This is part of our running series on cutting AI costs. See the full series → […]
AI on a Budget. This is part of our running series on cutting AI costs. See the full series → […]
OpenAI’s Astra produced machine-checked proofs for ten open maths problems. Impressive and verifiable, but there’s no product you can actually use.
OpenAI cut GPT-5.6 Luna 80% to about $0.20 per million tokens. The price of thinking dropped. Here’s what a solo operator should automate now.
Thinking Machines released Inkling-Small, an open-weights model that beats its bigger teacher on coding and reasoning. And you can run it yourself.
Alibaba launched Qwen3.8-Max on 3 August 2026: 2.4T parameters, a 1M-token window, open weights promised, and pricing well under Claude Opus 5.
DeepSeek retrained its open-weight V4 Flash to top its pricier Pro model on nine agent benchmarks, at $0.14 per million tokens. The catch is who scored it.
Google launched Gemini 3.6 Flash on 21 July, cheaper and leaner. But the 3.5 Pro flagship Pichai promised for June is still nowhere to be seen.
Moonshot released the full Kimi K3 weights on 27 July, the biggest open model yet. It tops WebDev Arena and undercuts closed US labs, if you can host it.
Reflection AI signed a $1B+ compute deal with Nebius through 2029, weeks after a reported $150M-a-month SpaceX deal. In modern AI, the moat is not the model. It is a confirmed delivery date for chips.
ZML just gave away a free tool that runs your AI models fast on any chip, not just Nvidia’s. It is a small, well-aimed kick at the CUDA lock-in that keeps the whole industry hostage.