Context caching
How repeated prompt prefixes are served from cache at a lower input price and how to read cache hits in the usage block.
When requests share a long prefix, the repeated part can be served from cache and billed at the cached input price shown in the catalog.
Many requests repeat the same opening: a long system prompt, a character card, a document you ask several questions about, or the earlier turns of a chat. Context caching serves that repeated opening from cache instead of processing it again. It is on for every request; there is nothing to enable and no extra parameter.
How it works#
- The cache matches the prefix of the prompt: the messages from the start of the list up to the first point where two requests differ.
- When a request starts with a prefix that a recent request already sent, the matching part is read from cache. The rest of the prompt is processed normally.
- The first request that sends a prefix fills the cache and is billed at the normal input price. Later requests with the same prefix get the cached price.
- Cache entries expire after a period without use, and short prompts are not cached. A hit is likely, not guaranteed, so do not build logic that depends on one.
Example: two requests with the same system prompt and document, but different questions.
text
Request 1: [system prompt][document][question A] -> fills the cache
Request 2: [system prompt][document][question B] -> [system prompt][document] read from cacheReading cache hits#
The usage block of every response shows how many prompt tokens came from cache:
JSON
"usage": {
"prompt_tokens": 3755,
"completion_tokens": 120,
"total_tokens": 3875,
"prompt_tokens_details": { "cached_tokens": 3584 }
}cached_tokens is part of prompt_tokens, not added to it. In this example, 3584 tokens are billed at the cached input price and the other 171 at the normal input price. In a stream, the same numbers are in the final usage chunk when stream_options.include_usage is set.
Pricing#
Cached input tokens cost less than regular input tokens. The cached price of each model:
| Model | Input (per 1M tokens) | Cached input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|---|
deepseek-roleplay | $0.50 | $0.10 | $1.00 |
lane-1 | $1.50 | $0.15 | $4.50 |
lane-1-pro | $4.50 | $0.45 | $7.50 |
A request is billed as: prompt tokens not read from cache at the input price, plus cached tokens at the cached input price, plus completion tokens at the output price.
Getting more cache hits#
- Put stable content first. System prompt, tool definitions, character cards and reference documents go at the start; the changing user question goes at the end.
- Keep the prefix byte-for-byte identical. A timestamp, a random id or a reordered tool list near the start changes the prefix and turns every request into a miss.
- Append, do not rewrite. In multi-round conversations, add new messages to the end of the list. Editing or summarizing old turns changes the prefix, so do it in larger steps rather than every round.
- Reuse the same model. Each model has its own cache; switching models starts from an empty cache.
- Send related requests close together. Questions about the same document hit more often when they are not spread far apart in time.