How Prompt Caching Cut Our AI Agent Costs by 63%

For years, every LLM call we made sent the same instructions, the same tool definitions, and the same growing chat history. We paid full price for all of it, every time. Then we shipped one release built around prompt caching.
In the 16 days before that release, 26% of our input tokens came from cache. In the 14 days after, 75% did. On a request that hits the cache, the median cached share is 93%, and the median request costs 79% less than it would have without caching. Across all traffic, our cost per million input tokens fell by 63%.
This is a case study for any team that runs LLM agents in production. It covers how agent costs compound, how prompt caching works inside the model, the four things we had to fix, the mistakes we found on the way, and 30 days of production data.
Where we started
We build AI agents for customer experience. Three results show the scale:
- Vodafone Qatar automated more than 400,000 conversations.
- International Medical Center automated 1 million customer service conversations.
- VM Group reduced support requests by 45%.
The full list is on our case studies page.
Through all of that, our attention went to what customers notice: answer quality, latency, and integrations. The cost of input tokens was not on that list. Model providers had offered prompt caching for a long time, and we had never examined it.
We started the work as an experiment and as good practice. It became one of the highest-return changes we have made to our agent runtime. Enabling the feature was the small part. Most of the work was in how we structured our prompts.
How agent costs compound
A single LLM call is cheap, and that hides the problem. An agent does not make single calls.
A chat agent runs in a loop. The end user sends a message. The model reasons, calls a tool, reads the result, reasons again, and answers. Then the end user replies, and the loop runs again.
Every one of those model calls sends everything that came before it again: the system prompt, the tool definitions, and the full conversation so far.
Here is how those calls add up on the bill, assuming cached tokens cost a quarter of the full input price.
The conversation produced only 6,500 tokens of real content. Without caching we pay for 31,500, because each call pays again for every call before it.
The total grows with the square of the conversation length. Long conversations are where an agent does the most useful work, and they are where cost rises fastest.
With caching, that same conversation bills the equivalent of 12,750 tokens, about 60% less. We still send the repeated part. It stops being expensive.
What prompt caching is and how it works
Before a model generates anything, it processes every token of the prompt and builds internal state for it. That computation is most of what you pay for on the input side. Prompt caching lets the provider keep that state for a few minutes.
If your next request begins with exactly the same tokens, the provider reuses the stored state for the matching part. It bills those tokens at a large discount.
Under the hood: the KV cache
A transformer handles a request in two phases. In the prefill phase, it reads the whole prompt in one pass. For every token, at every layer, it computes a key vector and a value vector. Those tensors are the KV cache.
In the decode phase, it generates output one token at a time, and each new token attends to the stored keys and values.
Prefill is the expensive part of a long prompt. Its compute grows with prompt length, and it must finish before the first output token appears. Prompt caching stores the KV tensors from one request so that a later request can skip prefill for the tokens they share.
This also explains why the cache can only match a prefix. Attention in these models is causal: the keys and values of a token depend on that token and on every token before it. Change one early token, and the tensors of all later tokens are different, even if their text is the same. There is nothing valid to reuse after the first difference.
Two practical effects follow:
- Lower latency as well as lower cost. A cache hit skips prefill for the shared part, so the first token arrives sooner on long prompts.
- Routing matters. The tensors live on specific machines. A request can only hit the cache if it reaches a machine that holds it.
The two properties that matter
- It is a prefix match. The cache covers the prompt from the first token up to the first difference, and nothing after that. One changed character near the top means everything after it is recomputed.
- It expires. Caches live for minutes, not days. They pay off on bursts of similar requests, which is exactly what a chat conversation is.
The first property is the one that matters for engineering. It means the order of your prompt is a cost decision.
Providers do not agree on the rules
We route model traffic through OpenRouter to OpenAI, Anthropic, and Google models. Each provider handles caching differently.
Every provider also has a minimum: prompts under roughly 1,024 tokens are never cached, whatever their shape.
These rules come from OpenRouter's prompt caching docs and OpenAI's prompt caching guide. Check them before you rely on a number, because they change with new models.
The Anthropic write premium matters. A breakpoint is a bet that the same prefix comes back before the cache expires. A cache read saves 90%, so one hit repays the write. A breakpoint on text that never repeats only adds 25% to the bill.
The four things we had to work on
Enabling caching was the small part. Real savings took four pieces of work, and we suggest the same order to any team.
1. See where the cost comes from
You cannot fix what you cannot see. "The LLM bill went up" is not something you can act on. We needed cost per request, per model, and per feature, with input tokens and output tokens separated.
The provider already reports all of that. Every response carries a usage block with the prompt tokens, the completion tokens, and the number of prompt tokens that came from cache. The ratio of cached tokens to prompt tokens is the one number to track, per model and per feature. If it sits near zero on a long, repetitive prompt, something near the top of that prompt changes on every request.
Two lessons from observing our own logs:
- Trust the cost the provider reports. A flat price table that charges every prompt token at the full rate ignores both halves of caching: the discounted reads and the more expensive writes. On a heavily cached request, that estimate can be several times too high. The provider's figure already has the arithmetic in it.
- Look at the prompt, not the template. The prompt a model receives is assembled from many parts, and the parts drift from what the template suggests. Reading a few real, fully assembled prompts, section by section, produced most of the findings below.
2. Layer the prompt from most stable to least stable
Because the cache is a prefix match, every prompt has an ideal order: the content that never changes first, the content that always changes last.
Less stable content that leaks into an earlier layer moves the cache boundary to that point. Our prompts grew without a plan over the years, and they had leaks.
This also became guidance for everyone on our team who writes prompts:
- Keep instructions, rules, and examples at the beginning.
- Put timestamps, end-user details, and request-specific context at the end.
- Do not reword the top of a prompt without a reason. A change there invalidates every cached token after it.
3. Treat tool definitions as part of the prompt
Tool definitions are easy to forget because they are not in the prompt text. They are tokens like any others, and providers serialize them ahead of the messages. When we measured a first-turn prompt for an agent with several tools, the tool descriptions were most of it.
That gives tool definitions three jobs.
They must be stable. Send the same tools, in the same order, with the same descriptions, on every call. If you build the tool list from a set or an unordered map, sort it first. A reordered list is a cache miss for the entire request.
They must carry their own meaning. Give a tool a precise name, a clear description, and a typed parameter schema. The model then understands how to use the tool from the definition itself. The alternative is to patch unclear tools with extra paragraphs in the system prompt. That makes the prompt longer, and it invites frequent edits, which break caching.
They must be as small as the model needs. Caching makes repeated tokens cheap, not free. We trimmed descriptions that repeated what the model already knew from elsewhere in the prompt. For agents with many tools, we also load tool definitions on demand: the model sees a small core set plus one lookup tool that can pull in the rest. That adds a lookup step, so it only pays off past a handful of tools, and a short tool list is better sent directly.
4. Move dynamic values out of the prefix
This is the most damaging kind of leak, and it is usually one line. Almost every agent needs the current time, so that it can handle a request like "book an appointment for tomorrow". The obvious place to put it is the system prompt:
<instructions> ... </instructions>
<current_time>2026-09-02T10:41:07.531924Z</current_time>
<conversation> ... </conversation>
<new_message> ... </new_message>A timestamp with sub-second precision is unique on every request. The match ends at that line. The conversation history behind it, the part that grows on every turn, can never be reused.
<instructions> ... </instructions>
<conversation> ... </conversation>
<new_message> ... </new_message>
<current_time>2026-09-02T10:41:07.531924Z</current_time>Reading our assembled prompts showed the same pattern in more than one place: the same request-specific values rendered by more than one section, each with its own clock reading. Three rules fixed it. Each value renders once, and one section owns it. Every section shares one clock reading per turn. Anything that is not plain text is left out. The model still sees the current time on every call, and the prefix in front of it stays identical.
Here is a short list of values to search your own prompts for: timestamps, request ids, session ids, end-user names, locale strings, feature flags, and A/B test variants.
Enabling caching across providers
With the prompts in order, the request itself needed two changes.
One breakpoint that works everywhere
Our first attempt used the top-level cache_control field, which asks OpenRouter to cache automatically. That field only works for Anthropic models, and our Gemini traffic showed no cache writes at all.
The fix was to put the system prompt in a content block with an explicit breakpoint:
{
"model": "<provider>/<model>",
"session_id": "<conversation id>",
"messages": [
{
"role": "system",
"content": [
{ "type": "text", "text": "...", "cache_control": { "type": "ephemeral" } }
]
},
{ "role": "user", "content": "... <dynamic values last>" }
],
...
}OpenRouter translates that one marker for each provider. For a supporting OpenAI model, it becomes a prompt_cache_breakpoint. For Anthropic and Google, it is a cache_control breakpoint. We have no provider-specific branches in our code, and Gemini started writing to its cache on the first test after the change.
We kept the five-minute TTL. The one-hour option doubles the write price. Our conversations are bursts of messages seconds apart, and every hit refreshes the window.
Sticky routing
The cache lives on specific machines, so a conversation must keep reaching the same backend. By default, OpenRouter derives a routing key from a hash of the opening messages. It only pins a backend after it observes a cache hit.
That leaves a gap. The turn that writes the cache can reach one backend, and the turn that would read it can reach another. The session_id field closes the gap. We send the conversation id, and sticky routing pins a backend from the first successful request.
What we saw in production
We exported 30 days of request logs from OpenRouter: about 55,000 requests on our production key, of which about 53,000 were LLM calls,, from 23 August to 21 September 2026. The caching release reached production on 8 September.
| Metric (LLM requests only) | Before: 23 Aug to 7 Sep | After: 8 Sep to 21 Sep |
|---|---|---|
| Share of input tokens served from cache | 26.4% | 74.5% |
| Requests over 1,024 tokens with a cache hit | 30.0% | 73.7% |
| Median cached share on a request that hits | 92.9% | 92.8% |
| Cache discount as a share of pre-discount cost | 4.3% | 64.1% |
| Cost per million input tokens | $0.86 | $0.32 |
The median cached share on a hit did not change, and that is the point. A request that hit the cache always had about 93% of its input served from it. The release did not make hits better. It made them happen: less than one third of eligible requests hit before, and nearly three quarters hit after.
The providers tell different stories:
| Provider | Cached share before | Cached share after | Cost per million input tokens |
|---|---|---|---|
| All providers | 26.4% | 74.5% | $0.86 to $0.32, 63% lower |
| OpenAI | 40.3% | 71.6% | $0.23 to $0.12, 47% lower |
| Anthropic | 0.0% | 84.0% | $2.24 to $0.69, 69% lower |
Anthropic is the clearest case. Caching is opt-in, we had never opted in, and the cached share was exactly zero. Nothing about those prompts was wrong. We had not asked.
OpenAI is the more instructive case. Caching was automatic, so 40% of input tokens already hit. The structure of our prompts and the routing gap wasted the rest. Moving dynamic values and pinning the backend raised the share to 72%, on top of caching that was already automatic.
Latency
We also compared time to first token after the release, on prompts of 4,000 tokens or more.
This is an observation, not a controlled test. The two groups contain different mixes of models and prompt sizes. The direction matches the theory: a cache hit skips prefill, and the first token arrives sooner.
What the data does not show
- Not every request can hit. Prompts under the provider minimum never cache. About one third of our requests are short calls to a small model, and they gain little.
- The first request of a conversation is a miss. So is any request that arrives after the cache expired.
- The model mix changed between the two periods. Cost per million tokens depends on which models serve the traffic. The per-provider results are a fairer comparison than the total.
- Thirty days is a short window. A longer window, with more model releases and traffic changes in it, would be a fairer test.
What we would tell another team
Measure input cost separately from output cost, and log cached tokens per request. If input dominates, caching pays. Then order every prompt from most stable to least stable, move every dynamic value to the end, keep tool definitions identical from call to call, and mark the static part with an explicit breakpoint rather than trusting each provider's automatic caching. Watch the cached share for a few days after the release.
None of this changed what our agents can do. It removed a cost we paid for the same computation on every turn of every conversation. If your prompts have a timestamp near the top, examine that first.
Build innovative AI Agents that deliver results
Recommended Reading: Check Out Our Favorite Blog Posts!

CodeKit: Code-Level Tools for Agents, Without the Wait

How We Built Media Retrieval for Tars Knowledge Bases





