AI glossary
Prompt caching
Reusing the processed beginning of a prompt across requests, which makes repeated long prompts cheaper and faster.
Many applications send the same long prefix again and again — a big system prompt, a document, a list of tools. With prompt caching the provider stores the processed version of that prefix for a short time; later requests that start with exactly the same content pay a much lower price for those tokens and get faster responses.
To benefit, keep the stable parts of your prompt at the start and put changing content (like the user’s question) at the end.
Example: A support assistant sends the same 50-page manual with every question. With caching, the first request processes the whole manual and later ones reuse that part. Anthropic says reading from the cache costs 10% of the normal input price on most of its models.
In practice
- Put the fixed parts first (instructions, documents, tools) and the variable parts last.
- Any change at the start of the prompt, even a date or a time, invalidates the cache.
- The cache lasts a limited time: it pays off when requests come in quick succession.

