Many LLM requests share a long, identical beginning: system instructions, tool definitions, documents or examples. Prompt caching lets providers reuse the processing of that prefix.
Benefits
- Lower cost for cached tokens, often substantially.
- Faster time to first token.
How It Works
The provider stores the processed state of a prompt prefix. Later requests with exactly the same prefix reuse it. Caches expire after a period without use.
Structuring Prompts for Caching
- Put stable content first: instructions, tools, reference documents.
- Put changing content last: the user's question, recent history.
- Keep the stable part byte-for-byte identical — even small changes break the cache.
- Avoid timestamps or random IDs early in prompts.
Good Candidates
- Chatbots with long system prompts.
- Question answering over the same large document.
- Agents resending tool definitions each step.
- Many-shot prompts with large example sets.
Provider Differences
Some providers cache automatically; others require marking cache points. Pricing and expiry vary — check documentation.
Measure
Track cache hit rates and savings.