How Prompt Caching Actually Cuts Your LLM API Bill
Prompt caching is one of the few AI cost levers that's almost entirely free once it's implemented: it costs engineering time to set up, and after that every cache hit is a request you're not paying full price for. The catch is that cache hit rate is fragile and drifts down quietly if nobody's watching it.
Here's how the saving actually works, worked through with a concrete example, and what tends to erode it after launch.
How caching actually reduces the bill
Most model providers charge a lower rate for tokens that were already processed in a recent, matching request, typically the shared prefix of a prompt: system instructions, a knowledge base excerpt, or conversation history repeated across turns. A cache hit means that shared portion is billed at the reduced rate instead of the full rate. The saving scales with how much of your prompt is repeated content versus genuinely new content on each call.
Where cache hit rate actually comes from
Hit rate depends on how much of your prompt structure is stable across requests and how close together in time those requests arrive, since most providers only cache for a limited window. A support chatbot with a long, static system prompt and frequent messages in the same conversation will see a high hit rate. A one-off request with a unique prompt every time will see close to none, no matter how the caching is configured.
A worked example
Say your support bot sends a 2,000 token system prompt plus a short user question on every call, and conversations average six turns. On the first turn, the whole prompt is new and billed at full rate. On the next five turns, the system prompt portion, most of the token count, can be served from cache at the reduced rate, while only the new turn's content is billed in full. Across a full conversation, the blended cost per turn drops substantially once you look past the first message, purely from restructuring the prompt so the stable part comes first and stays byte for byte identical between calls.
The same logic applies to anything else that repeats across calls: a knowledge base excerpt injected into every request, a fixed set of tool definitions, or a style guide the model is told to follow. Anything static that currently gets reassembled and sent fresh on every call is a candidate to move to the front of the prompt and keep identical, purely to make it eligible for the cached rate.
What quietly breaks it in production
Cache hit rate erodes for reasons that rarely show up in a code review:
- A timestamp, request ID, or other unique value inserted anywhere in the cached portion of the prompt
- Reordering or lightly editing the system prompt during a routine update, which changes it just enough to miss the cache
- Conversation gaps longer than the provider's cache window, common in support tools where a customer replies hours later
- Running the same logical prompt through slightly different formatting in different parts of the codebase
Tracking it as a number finance watches
Cache hit rate should sit next to token spend on whatever dashboard tracks AI cost, not live only in an engineering dashboard nobody outside the team checks. A quiet drop in hit rate looks identical to a rise in traffic on the invoice, and the two need very different responses: one is a cost containment fix, the other is a usage growth story. Without hit rate visible next to spend, finance has no way to tell which one actually happened.
When a low hit rate is the wrong thing to chase
Not every workload can reach a high cache hit rate no matter how well the prompt is structured, and treating a stubbornly low number as a problem to fix can waste engineering time chasing a target the workload was never going to hit. A one-off research query, a document summarization job where every document is different, or a batch classification run over unique records all have little or no repeated content to cache in the first place.
For those workloads, the lever that actually moves cost sits somewhere else: a smaller or cheaper model for the task, a shorter prompt, or batching several similar requests into fewer calls if the use case allows it. Spend the analysis time figuring out which lever fits the workload rather than assuming caching applies everywhere token cost shows up.
The useful question isn't why a hit rate is low in isolation, it's whether the workload has a structurally cacheable shape at all. A support chatbot with a long static system prompt clearly does. A one-off research assistant answering a different question every time from a different context clearly doesn't, and no amount of prompt restructuring will change that.
What Good Looks Like
Good looks like a stable or improving cache hit rate tracked alongside token spend, with a known owner when it drops.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Why did our cache hit rate drop after we added a new feature?
The most common cause is a prompt structure change: adding a timestamp, a random identifier, or reordered content near the top of the prompt breaks the exact match caching depends on. Check whether anything in the cached portion of the prompt varies between calls that used to be identical.
Does prompt caching affect response quality?
No, caching affects billing and latency for the repeated portion of the prompt, not the content or quality of the model's response. It's a cost and speed optimization, not a tradeoff against accuracy.
How long should we expect a cache to stay warm?
This varies by provider and is usually on the order of minutes, not hours, so caching helps most within an active session and much less for requests spread far apart in time. Check your specific provider's documentation for the current window, since providers do change this.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Open Source vs Proprietary LLMs: What Each Choice Costs Your Margin
A CFO's side-by-side look at what open source and proprietary language models actually cost once you include hosting, tuning and engineering time.
When You Can Capitalize LLM Fine-Tuning Costs Under ASC 350-40
How ASC 350-40's three development stages apply to LLM fine-tuning and RAG pipeline work, so you know which costs to expense and which to capitalize.
Reselling a Third-Party API: Protecting Your Margin
How to structure contract terms and pricing so a vendor's price increase or rate limit change doesn't quietly erase the margin you're reselling their API on.
Negotiating a Cloud Minimum Spend Commitment Without Overcommitting
How hyperscaler minimum spend commitments are structured, what happens to credits you don't use, and the terms worth pushing back on before you sign.
Prompt Engineering Time vs Fine-Tuning: Where the ROI Actually Breaks
Why prompt engineering hours quietly become the more expensive option over time, and how to tell when fine-tuning would actually cost less.
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.