Google will let you call Gemini 3.8 Flash, its best Flash model, for nothing. The same model costs $0.75 per million input tokens and $3.75 per million output tokens on the paid tier, and those prices double on 1 January 2027. Meanwhile Cerebras, which spent much of this year advertising a million free tokens a day, now asks for a card and hands you $5 of credit that dies after 30 days.
That is the free AI market in September 2026 in two sentences: still generous in places, shrinking in others, and full of small print that decides whether your side project runs for free or runs up a bill on a dashboard nobody checks. This instalment of our AI on a Budget series maps which free tiers are worth your time, which open models are worth paying pennies for, and where each one breaks.
How we got here
Free tiers were a land grab. In 2024 and 2025 every inference provider wanted developers hooked on its API, so free quotas were loose and nobody asked many questions. Two things changed. Open-weight models got good enough that people started running real products on free quotas, and reasoning models started burning tens of thousands of tokens per answer. Free capacity got expensive to give away.
The response has been consistent across providers: cap tokens per day as well as requests, limit quotas per organisation or project rather than per API key, and push the best models behind a paywall. Google’s newest Pro model, Gemini 3.1 Pro Preview, is listed as “Not available” on the free tier. Cerebras dropped its permanent free allowance entirely. Nobody is being villainous here. Free is a sales channel, and the channel is being tightened.
The current map: four free tiers worth knowing
Google Gemini (AI Studio)
Still the most generous by model quality. Gemini 3.8 Flash, 3.5 Flash-Lite and even the older 2.5 Pro are listed as free of charge for input and output. Limits are applied per Google Cloud project, not per key, and daily quotas reset at midnight Pacific time. Google no longer publishes a neat per-model table; your live limits sit on the AI Studio rate-limit page.
You pay with data. Google’s own pricing page states that free-tier content is “used to improve our products”, while paid-tier content is not. Fine for a hobby chatbot. Not fine for client contracts, medical notes or anything under an NDA. Free tier also loses Google Search grounding and context caching.
Groq
Groq’s free plan covers open models at very high speed. Current published free limits for openai/gpt-oss-120b, gpt-oss-20b and qwen/qwen3.8-27b are 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute and 200,000 tokens per day. Limits apply at organisation level, so spinning up extra keys gains you nothing. Cached tokens do not count towards the limits, which rewards anyone who read our prompt caching piece.
Tokens per day will bite first. A workload of 300 requests at roughly 2,500 tokens each is 750,000 tokens a day, nearly four times the free ceiling. Realistically, Groq free covers about 80 of those calls daily.
OpenRouter free models
OpenRouter exposes a rotating set of models with IDs ending in :free. Every one is capped at 20 requests per minute. Daily caps depend on whether you have ever bought credits: 50 requests a day if you have purchased less than $10 in total, 1,000 a day once you have bought $10 or more, and that higher ceiling sticks even if your balance drops to zero. A one-off $10 top-up is the best value purchase in this whole article.
Free variants are served by whichever provider has spare capacity, availability shifts, and the list of free models changes without much warning, so reliability suffers. Build a fallback to a paid model or accept that your app will sometimes return a 429.
Cerebras
Cerebras has ended its free allowance and replaced it with a trial. New accounts get $5 in credits after adding a verified payment method, the credits expire after 30 days, and the Free Trial limits for gpt-oss-120b and qwen-3.8-27b are 5 requests per minute and 1 million tokens per day. Cerebras states plainly that there is no permanently free tier. Useful for benchmarking its speed; useless as a long-term home.
Open models worth paying pennies for
Once you outgrow free quotas, open-weight models on third-party hosts are usually the cheapest step up. Three are worth shortlisting.
- DeepSeek V4 Flash. 284B parameters with 13B active, 1M-token context, MIT licence, released 24 April 2026. On OpenRouter most hosts charge around $0.14 input and $0.28 output per million tokens, with the cheapest discounted listing at $0.0399 and $0.0798. Artificial Analysis scores its max-effort mode at 89.4% on GPQA Diamond.
- gpt-oss-120b. OpenAI’s open-weight model, available on both Groq and Cerebras free allowances. A sensible default when you want one model you can run free today and pay for tomorrow without rewriting prompts.
- Mistral Small 4. 119B parameters with only 6.5B active, Apache 2.0. Roughly 111GB of VRAM at FP8, around 75GB with community 4-bit builds. The European option if data residency matters.
Skip the headline giants if you are on a budget. Kimi K3 is 2.8 trillion parameters and its weights alone come to about 1.56TB. GLM-5.2 needs roughly 245GB of VRAM even at 2-bit quantisation. Impressive on leaderboards, pointless on a hobby budget unless someone else hosts them.
What it costs once free runs out
Take that same 300-calls-a-day workload: 2,000 input tokens and 500 output tokens per call, so 600,000 input and 150,000 output tokens daily.
- Gemini 3.8 Flash, paid: about $1.01 a day, roughly $30 a month at 2026 prices, rising to about $60 from January.
- DeepSeek V4 Flash at typical $0.14/$0.28: about $0.13 a day, under $4 a month.
- DeepSeek V4 Flash at the cheapest listing: about $0.04 a day, around $1 a month.
For a lot of internal tools, the gap between “free” and “cheap open model” is the price of a coffee each month. The gap between free and a proprietary paid tier is closer to a streaming subscription, and it is about to double.
A sensible free-first setup
- Prototype on Gemini free for quality, with nothing confidential in the prompts.
- Put $10 into OpenRouter once to unlock 1,000 free requests a day and a single API for fallbacks.
- Use Groq free for fast, short calls where 200,000 tokens a day is enough.
- Route anything sensitive or high-volume to a paid open model such as DeepSeek V4 Flash, and set a hard spend cap.
- Pair it with model routing so the free tier handles easy requests and paid capacity handles the rest.
What this means
- Free tiers are real but narrowing: tighter daily token caps, per-organisation limits, and top models removed.
- Google’s free Gemini is the best quality you can get for nothing, paid for with your data.
- OpenRouter’s $10 one-off purchase is the single best value move for hobbyists.
- Cerebras is a 30-day trial now, so do not build on it.
- Past free, open models like DeepSeek V4 Flash run typical small workloads for under $4 a month.
- Gemini 3.8 Flash paid prices double on 1 January 2027. Budget for it now or plan a switch.
Next in the series: quantisation, properly explained, so you can fit these open models onto hardware you already own.
Did you know: a peer-reviewed Nature paper put the reasoning training for DeepSeek R1 at roughly $294,000, built on a base model that cost around $6 million.
Sources
- Google: Gemini Developer API pricing
- Google: Gemini API rate limits
- Groq: Rate limits
- OpenRouter: Limits
- Cerebras: Rate limits and Free Trial FAQ
- OpenRouter: DeepSeek V4 Flash pricing and benchmarks
- Thunder Compute: Best open source LLMs (September 2026)
Related on Top Tool Stack: AI on a Budget: cut your AI costs · Prompt caching