Prompt Caching: Making Repetition Cheap
Prompt caching can reduce the cost and latency of processing a repeated request prefix, but it does not shrink context, select relevant information, create memory, or improve bad evidence.

A documentation assistant sends the same tool definitions, instructions, and product manual on many requests. Only the session state and latest question change.
Without caching, the provider may have to process that repeated prefix again for every model call. With an eligible prompt cache, a later request can reuse provider-maintained intermediate processing for an identical earlier portion of the prompt.
That can make repeated input cheaper or faster under the provider’s current rules. It does not remove the input from the model’s context, decide whether the manual is relevant, or preserve the conversation after the cache expires.
Two calls / one repeated prefix
Caching changes repeated processing economics- Tool definitions
- System instructions
- Static reference material
- Session state
- Current task
- Same tool definitions
- Same system instructions
- Same static reference material
- Updated session state
- New current task
The reusable object is a prefix
A prompt cache generally operates on a prefix: the ordered content from the start of a rendered request to a provider-defined or application-marked boundary.
A normalized request might be assembled as:
stable tool definitions
-> stable system instructions
-> stable reference material
-> changing session state
-> current user task
On the first eligible request, the provider processes the prefix and creates or writes a cache entry. On a later request, the provider looks for a matching prefix under its cache policy. A hit allows the service to reuse the cached representation and process the changing tail normally.
“Normally” matters. The model still generates a fresh response. Prompt caching is not a database of prior answers and does not force deterministic output. It optimizes repeated input processing.
Why ordering changes reuse
Compare two layouts containing the same broad information.
Volatile first
current question
-> timestamp
-> changing session summary
-> tools
-> instructions
-> static manual
Stable first
tools
-> instructions
-> static manual
-> changing session summary
-> current question
In the first layout, the request differs almost immediately. The later static material is no longer part of one long shared prefix. In the second, repeated material forms a stable early region and volatile content follows it.
This does not mean stable-first is a universal semantic instruction order. Providers render roles, tool schemas, images, and messages differently, and instruction authority must be preserved. It means cache-aware request assembly should keep material that is both reusable and validly ordered stable before the changing suffix when the target API supports that arrangement.
Even tiny differences can matter under exact-match policies:
- inserting a current timestamp into an otherwise stable instruction;
- reordering tool definitions;
- changing whitespace or serialized field order where the provider includes it;
- updating one reference passage;
- changing model or reasoning configuration;
- placing user-specific state before the intended boundary.
The exact invalidation rules are API-specific. Measure hits from provider usage records rather than assuming that text which looks similar to a human is cache-equivalent.
Write, read, expire, miss
Cache economics have a lifecycle.
- Write or creation: The first eligible request processes content and establishes reusable cache state. Some providers price this differently from ordinary input.
- Read or hit: A later request reuses an eligible matching prefix. The provider may report cached input separately and may apply different input pricing or latency behavior.
- Refresh: A successful reuse may extend the entry’s lifetime under the provider policy.
- Expiration or eviction: A time-to-live, inactivity period, capacity policy, deployment detail, or configuration can make the prior entry unavailable.
- Miss: The request no longer matches or the entry is unavailable, so the service processes the input without that reuse and may create new cache state.
Anthropic and OpenAI both currently document prefix-oriented caching, but their field names, breakpoint behavior, minimum eligible lengths, cache lifetimes, matching details, model support, and write/read pricing are not the same. Those details also change over time. Consult the target model and API documentation, then inspect reported cache reads and writes in real traffic.
No vendor-neutral lesson should promise one discount, TTL, minimum length, or automatic behavior.
Caching and compaction solve different problems
Caching can make repeated context less expensive to process. Compaction changes the representation so less or different context is carried forward.
| Mechanism | Primary question | What changes |
|---|---|---|
| Prompt caching | Can this repeated prefix be processed more efficiently? | Provider economics or latency for eligible repeated input |
| Compaction | What information should remain in active context? | Context size, representation, and fidelity |
If an application caches a 70,000-token prefix, that prefix can still occupy 70,000 tokens of the call’s context accounting. A cheaper large prefix is not a smaller prefix. It can still leave less output headroom, include stale evidence, bury a current instruction, or exceed a model limit.
Likewise, compacted state is not automatically cache-friendly. If the summary changes on every turn and appears early, it creates a volatile prefix. A system may therefore combine the mechanisms:
cache stable tools and instructions
-> compact changing session state
-> retrieve exact evidence for the current task
-> append the latest input
Each arrow has a separate purpose.
Caching is not memory
A cache can survive briefly between requests, but that does not make it persistent user memory.
Persistent memory involves an external system retaining information so it can be selected for a future interaction. Prompt caching retains provider-specific reusable processing for eligible matching content under a bounded cache policy.
The contrast becomes obvious after expiry:
stored user preference
-> application retrieves it tomorrow
-> preference enters current context
cached prompt prefix
-> cache expires or no longer matches
-> request receives no cache reuse
The first path can restore information because a durable store and retrieval policy exist. The second path only changes how efficiently an already supplied prefix is processed. It does not decide to add a forgotten preference back into the request.
Caching is not retrieval
Retrieval starts from an information need and fetches candidate evidence. Prompt caching starts from repeated request content and tries to reuse its processing.
A RAG system might retrieve a different policy passage for each question, making that region volatile. Or it might repeatedly send a large static handbook that happens to be cacheable. The retrieval and caching decisions can coexist, but one does not imply the other.
Most importantly, a cache hit says nothing about source quality. A stale document can be cached perfectly. So can an irrelevant tool catalog, a contradictory instruction, or a summary that lost the decisive exception.
Tools create a caching and selection problem
Tool definitions can consume context before any tool runs. A system with many tools may repeatedly send names, descriptions, schemas, and usage instructions on every turn.
If the tool surface is stable and the provider supports caching it, those definitions may form part of a reusable prefix. Yet caching the catalog does not make every tool relevant. Exposing unnecessary tools can still increase context size, complicate selection, create permission risk, and add ambiguity.
Two policies remain separate:
- tool routing or selection decides which tools this task should expose;
- prompt caching may reduce repeated processing for the selected stable definitions.
“Cache all tools” is not a substitute for choosing the appropriate tool boundary.
Design for measured reuse
A cache-aware request pipeline can ask:
- Which content is genuinely shared across calls?
- Which shared content belongs early under the provider’s valid request structure?
- Where does volatile user, session, tool-result, or retrieved state begin?
- Which changes invalidate reuse under this model and API?
- What are the current write, read, lifetime, privacy, and pricing rules?
- Do usage records show enough reads to justify the writes and assembly constraints?
Then it must ask a separate set of context-quality questions: Is the repeated information still needed? Is it fresh? Is it authoritative? Does it crowd out more useful evidence?
Prompt caching makes repetition cheaper when the provider can recognize a stable prefix. It does not make repetition wise. Context selection and compaction still determine what the model receives; verification still determines whether the result is supported.
References
- Prompt cachingAnthropic
- Prompt cachingOpenAI
- Conversation stateOpenAI