A 30-turn support conversation on Claude Sonnet 4.6 costs about 84 cents in input tokens if you resend the full transcript on every turn. Trim it properly and the same conversation costs about 32 cents. At 10,000 conversations a month that gap is roughly $5,200, and closing it takes something like forty lines of code.
The bill grows the way it does because your transcript is billed as input on every single call. Turn 30 re-reads turns 1 through 29 and pays full price for all of them again. A conversation that feels linear to the user costs you a triangular number of tokens, and most teams never spot it, because the per-call figure looks trivial right up until the monthly invoice stops looking trivial.
This is the seventh instalment of AI on a Budget, our series on cutting AI spend without downgrading the output. Previous entries covered prompt caching, model routing, local models, batch APIs, semantic caching and draft-then-refine. This one is about the tokens you never send at all, which are the cheapest tokens available anywhere.
The maths nobody runs
Take a support bot averaging 600 tokens per turn, user message and model reply combined. Send the whole history each time and by turn 30 you have billed 600 x (1+2+…+30), or 279,000 input tokens, for one conversation. On Claude Sonnet 4.6 at $3 per million input tokens that is $0.84. On Opus 4.8 at $5 it is $1.40. On GPT-5.6 Terra at $2 it is $0.56, and on the cheap GPT-5.6 Luna tier at $0.20 it is about 6 cents.
Now cap the window at the last six turns and carry a 300-token rolling summary of everything older. Turns one to six cost the same 12,600 tokens. Every turn after that sends roughly 3,900. Total for the conversation: about 106,000 tokens. That is a 62% cut, and the model still knows who the customer is and what they asked for.
The summary is not free. Regenerating it four times per conversation on Claude Haiku 4.5 at $1 in and $5 out costs about 2.2 cents, or $220 a month across 10,000 conversations. Net saving on the Sonnet numbers: a shade under $5,000 a month, for a workload most teams would describe as small.
Four moves, in order of effort
- Sliding window. Keep the system prompt, the last N turns, and nothing else. N of 6 to 10 covers most support and chat work. Highest saving per hour of engineering in this entire series, prompt caching aside.
- Rolling summary. Every N turns, have a cheap model compress everything outside the window into 200 to 400 tokens: who the user is, what they want, what has been agreed, what is still open. Use Haiku 4.5 or GPT-5.6 Luna for this. Never the model doing the real work.
- Clear dead tool results. In agent workloads the transcript is mostly file contents and search results the model already read and acted on. They are the fattest and most disposable blocks in the context.
- Stop sending what you could fetch. If a 4,000-token product catalogue rides along in every request and gets used in one conversation out of twenty, make it a tool call instead.
The providers will now do most of this for you
Anthropic’s context editing beta, enabled with the header context-management-2025-06-27, strips stale tool results server-side. The clear_tool_uses_20250919 strategy triggers by default at 100,000 input tokens, keeps the three most recent tool use and result pairs, and swaps the rest for a placeholder. Anthropic’s own token-counting example shows a request falling from 70,000 tokens to 25,000, a 64% cut on a single call. Your client keeps the full history; only what reaches the model gets edited.
Server-side compaction goes further. Added in January 2026 as compact_20260112 behind the compact-2026-01-12 beta header, it summarises the conversation past a threshold you set, minimum 50,000 tokens and 150,000 by default, then continues from the summary. Anthropic’s documentation shows histories approaching 100,000 tokens collapsing back to 2,000 or 3,000. Worth noting that Claude Code’s interactive auto-compact fires at roughly 98% of the effective window, which is far too late if cost rather than overflow is your problem. Set your own threshold much lower.
OpenAI’s Responses API is where people talk themselves into a saving that does not exist. Passing previous_response_id means you only transmit the new user message, and the server-side state storage itself is free. You are still billed for every token of the reconstructed history on every call. It makes your code simpler. It does nothing to your bill, and it hides the growth curve instead of flattening it.
Where this goes wrong
Trimming fights prompt caching, and caching usually wins. Cached input bills at 10% of standard rates on OpenAI and 90% off on Anthropic, but the cache matches on a stable prefix. Cut something out of the middle of your history and you invalidate everything after it, then pay cache-write costs to rebuild. Anthropic’s answer is the clear_at_least parameter: refuse to clear at all unless the clear is large enough to justify breaking the cache. Set it to 5,000 tokens or higher. Trim in chunks, on a trigger, never continuously.
The second failure is recall, and the research here cuts in your favour. Chroma’s context-rot report tested 18 frontier models including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, holding task difficulty constant and varying only input length. Every model degraded as the input grew, even on trivial copy-and-retrieve tasks. Shorter context is frequently more accurate as well as cheaper. Summarisation still loses things, though, and what it loses is usually account numbers, exact dates and constraints agreed twenty turns ago. Pin those. Extract the hard facts into a small structured state object you send verbatim every time, and let the summary carry the narrative.
Third and most embarrassing: do not summarise with your main model. Compressing 4,000 tokens on Opus 4.8 costs five times what Haiku 4.5 charges and produces a summary nobody will ever read.
What to actually do this week
- Log
input_tokensper call and plot it against turn number. If that line slopes upward, you are paying the triangle. - Cap the window at eight turns. Add a 300-token rolling summary generated on your cheapest model. Pin IDs, dates and dollar figures verbatim outside the summary.
- On Anthropic, switch on
clear_tool_uses_20250919withtriggerat 30,000 input tokens andclear_at_leastat 5,000 for any agent that reads files or searches the web. - Preview the saving before you change production behaviour. The token-counting endpoint accepts the same
context_managementblock and reportsoriginal_input_tokensalongside the edited count.
What this means
- Typical saving: 50 to 70% of input spend on multi-turn workloads, stacking on top of caching and routing rather than competing with them.
- Effort: a day for the window and rolling summary, an afternoon for the Anthropic flags.
- Who it pays off for: anyone running chat, support or agents past about ten turns. Single-shot classification and extraction jobs get nothing from this, so do not bother.
- The risk to manage: cache invalidation and lost specifics. Trim in large chunks, pin the facts you cannot afford to lose.
- The number to watch: average input tokens per call. If it is climbing month over month while your user count is flat, context discipline is the cheapest fix on your list.
Did you know: the numbers 1 to 30 add up to 465, which is why thirty chat turns billed with full history cost more than fifteen times what those same thirty turns cost with no history attached, and why an AI bill can grow far faster than the user count while nothing about the product changes at all.
Sources
- Anthropic, Context editing (Claude Platform Docs)
- Anthropic, Compaction (Claude Platform Docs)
- Context rot: how long inputs degrade LLM accuracy (on Chroma’s 18-model study)
- OpenAI Responses API: the cost of state retention
- CloudZero, Claude pricing in 2026
- CloudZero, OpenAI API pricing in 2026
Related on Top Tool Stack: AI on a Budget: cut your AI costs · Prompt caching