Pay full price once. Reuse at the cache rate.
Repeated prompt tokens bill at the provider’s cache rate — often a fraction of the input price — and the saving lands on the receipt rather than in a projection.
- Cache-eligible prefixes are billed at the provider’s cache rate across 5-minute and 1-hour ephemeral windows.
- cached_tokens appears on every receipt, so the saving is auditable rather than asserted.
- Nothing to configure in your client — cache_control is passed straight through to the provider that supports it.
Most of your prompt never changes.
The system prompt, the schema, the few-shot examples, the document you are asking about — the same tokens, sent again and again. Cached, they bill at the provider’s cache rate, which is routinely a tenth of the input price.
And you can check the arithmetic.
cached_tokens is on every receipt next to the tokens you paid full price for. The saving is a number you can add up at the end of the month, not a percentage in a pitch.