Skip to content

Prompt Caching

How caching the repeated parts of prompts cuts cost and latency for LLM applications, and how to structure prompts to benefit.

Editorial team 1 min read

Many LLM requests share a long, identical beginning: system instructions, tool definitions, documents or examples. Prompt caching lets providers reuse the processing of that prefix.

Benefits

  • Lower cost for cached tokens, often substantially.
  • Faster time to first token.

How It Works

The provider stores the processed state of a prompt prefix. Later requests with exactly the same prefix reuse it. Caches expire after a period without use.

Structuring Prompts for Caching

  • Put stable content first: instructions, tools, reference documents.
  • Put changing content last: the user's question, recent history.
  • Keep the stable part byte-for-byte identical — even small changes break the cache.
  • Avoid timestamps or random IDs early in prompts.

Good Candidates

  • Chatbots with long system prompts.
  • Question answering over the same large document.
  • Agents resending tool definitions each step.
  • Many-shot prompts with large example sets.

Provider Differences

Some providers cache automatically; others require marking cache points. Pricing and expiry vary — check documentation.

Measure

Track cache hit rates and savings.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026