Context caching

How repeated prompt prefixes are served from cache at a lower input price and how to read cache hits in the usage block.

When requests share a long prefix, the repeated part can be served from cache and billed at the cached input price shown in the catalog.

Many requests repeat the same opening: a long system prompt, a character card, a document you ask several questions about, or the earlier turns of a chat. Context caching serves that repeated opening from cache instead of processing it again. It is on for every request; there is nothing to enable and no extra parameter.

How it works#

  • The cache matches the prefix of the prompt: the messages from the start of the list up to the first point where two requests differ.
  • When a request starts with a prefix that a recent request already sent, the matching part is read from cache. The rest of the prompt is processed normally.
  • The first request that sends a prefix fills the cache and is billed at the normal input price. Later requests with the same prefix get the cached price.
  • Cache entries expire after a period without use, and short prompts are not cached. A hit is likely, not guaranteed, so do not build logic that depends on one.

Example: two requests with the same system prompt and document, but different questions.

text

Request 1: [system prompt][document][question A]   -> fills the cache
Request 2: [system prompt][document][question B]   -> [system prompt][document] read from cache

Reading cache hits#

The usage block of every response shows how many prompt tokens came from cache:

JSON

"usage": {
  "prompt_tokens": 3755,
  "completion_tokens": 120,
  "total_tokens": 3875,
  "prompt_tokens_details": { "cached_tokens": 3584 }
}

cached_tokens is part of prompt_tokens, not added to it. In this example, 3584 tokens are billed at the cached input price and the other 171 at the normal input price. In a stream, the same numbers are in the final usage chunk when stream_options.include_usage is set.

Pricing#

Cached input tokens cost less than regular input tokens. The cached price of each model:

ModelInput (per 1M tokens)Cached input (per 1M tokens)Output (per 1M tokens)
deepseek-roleplay$0.50$0.10$1.00
lane-1$1.50$0.15$4.50
lane-1-pro$4.50$0.45$7.50

A request is billed as: prompt tokens not read from cache at the input price, plus cached tokens at the cached input price, plus completion tokens at the output price.

Getting more cache hits#

  • Put stable content first. System prompt, tool definitions, character cards and reference documents go at the start; the changing user question goes at the end.
  • Keep the prefix byte-for-byte identical. A timestamp, a random id or a reordered tool list near the start changes the prefix and turns every request into a miss.
  • Append, do not rewrite. In multi-round conversations, add new messages to the end of the list. Editing or summarizing old turns changes the prefix, so do it in larger steps rather than every round.
  • Reuse the same model. Each model has its own cache; switching models starts from an empty cache.
  • Send related requests close together. Questions about the same document hit more often when they are not spread far apart in time.