AI is cheap to start with and expensive to scale. The gap between a hobby project and a real bill is enormous, and most of it is waste: paying frontier prices for trivial jobs, re-sending the same context on every call, renting compute you could own. This series is a running, practical guide to closing that gap. Real numbers, exact settings, honest trade-offs, no hype.
Every instalment stands alone, but together they stack. Pull two or three of these levers at once and a bill that reads like a mortgage payment starts reading like a phone plan.
The series so far
Prompt Caching: the boring trick that cuts your LLM bill by 90%
Stop paying to re-read the same system prompt on every call. Cache the stable prefix and cache reads drop to roughly a tenth of the price. The single highest-return hour on this list.
LLM Routing: stop paying frontier prices for easy questions
Most of your queries are trivial. A router sends the easy majority to a cheap model and only escalates the hard ones, cutting costs 40 to 85% at near-identical quality.
Running Open Models Locally: kill your API bill in 2026
A used graphics card and a quantised open model can replace a chunk of your API spend outright. The real break-even maths, the setup, and where local still loses.
Batch APIs: half price on every AI job that can wait until tomorrow
OpenAI, Anthropic and Google all take a flat 50% off input and output for deferrable work. The real limits, how it stacks with caching, and the operational cost you pay in exchange.
Semantic Caching: reuse LLM answers and cut repetitive API spend by half
Embed each question, match it to one you have already answered, and skip the model call. 30 to 50% off conversational traffic, with the threshold and false-positive maths you need to run it without serving wrong answers.
The draft-then-refine pattern: let the cheap model write and the expensive one edit
A cheap model writes the first pass, a frontier model returns patches instead of a rewrite. 35 to 55% off generation-heavy workloads, with the worked maths and the lazy version that costs you 20% more than skipping the draft.
Context discipline: the cheapest tokens are the ones you never send
Every turn re-bills the whole transcript, so a 30-turn chat costs a triangular number of tokens. Sliding windows, rolling summaries and server-side context editing take 50 to 70% off input spend, with the cache-invalidation trap explained.
Free AI still exists. Mind the small print.
Gemini, Groq, OpenRouter and Cerebras free tiers mapped with their real limits, plus the cheap open models to use once free runs out. A $10 one-off top-up is the best value move in the piece.
Quantisation explained: GGUF, AWQ and GPTQ without the jargon
Llama 3.1 8B drops from 14.96 GiB to 4.58 GiB at Q4_K_M and generates about 2.5 times faster. Which format fits a laptop, which fits a GPU server, and why the serving kernel matters more than the algorithm.
Coming up in the series
The plan from here, roughly in the order they save you the most for the least effort:
- Structured outputs and prompt hygiene: smaller, tighter prompts and typed responses that stop you paying for waffle.
- RAG and embeddings on a budget: retrieval that answers from cheap storage instead of expensive context.
New instalments land regularly. The fastest way to catch each one is the newsletter below.
Did you know: the cheapest possible API call is the one you never make. Caching, batching and routing all exist for the same reason: computers hate doing the same expensive work twice, and so should your invoice.
