How it works
Claude: you mark where the fixed part of the prompt ends. The first call writes it to cache (1.25× the input price, or 2× for a 1-hour TTL); later calls read it for a fraction of the price: 2.5% to 10% of input, depending on the model. The cache expires after 5 minutes unused (or 1 hour).
OpenAI: caching is automatic for repeated prefixes and bills cached tokens at 10% of input, with no write cost.
When it pays off
When the fixed part is large and calls come often enough that the cache doesn't expire. With a low hit rate on Claude, each write costs more than the reads save.
Put what never changes (instructions, tools, documents) at the start of the prompt and what changes per call at the end.