Tokens & token usage

What a token is, how prompt, cached and completion tokens are counted, and where to read the exact usage of every request in the API response.

Models read and write text as tokens. Every response carries a usage block with the exact prompt, cached and completion token counts the request was billed for.

A token is the unit a model reads and writes: a short word, part of a longer word, a number or a punctuation mark. As a rough guide, one English word is about 1.3 tokens and 1,000 tokens are about 750 words. Each model has its own tokenizer, so the same text can give different counts on different models. The only exact count is the one the API returns.

The usage block#

Every chat completion response ends with a usage object. These are the numbers your request is billed for.

JSON

"usage": {
  "prompt_tokens": 1250,
  "completion_tokens": 340,
  "total_tokens": 1590,
  "prompt_tokens_details": { "cached_tokens": 1024 },
  "completion_tokens_details": { "reasoning_tokens": 120 }
}
ParamValue
prompt_tokensTokens in the input: system prompt, messages, tool definitions and images.
completion_tokensTokens the model generated, including reasoning tokens.
total_tokensprompt_tokens + completion_tokens.
prompt_tokens_details.cached_tokensThe part of prompt_tokens served from cache and billed at the cached input price. 0 when nothing was cached.
completion_tokens_details.reasoning_tokensThe part of completion_tokens the model spent thinking before the answer. Billed at the output price, not on top of it.

The cost formula and the prices are on Models & pricing.

Usage in streaming responses#

A streamed response does not include usage unless you ask for it. Set stream_options.include_usage to true and the last chunk before data: [DONE] carries the usage object with an empty choices array.

Python

stream = client.chat.completions.create(
    model="lane-1",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
    elif chunk.usage:
        print("\n", chunk.usage)

Node.js

const stream = await client.chat.completions.create({
  model: "lane-1",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,
  stream_options: { include_usage: true },
})

for await (const chunk of stream) {
  if (chunk.choices.length > 0) process.stdout.write(chunk.choices[0].delta.content ?? "")
  else if (chunk.usage) console.log("\n", chunk.usage)
}

The request is billed the same way whether or not you ask for the usage chunk.

Reasoning tokens#

Some models think before they answer. The thinking text arrives in the reasoning_content field of the message (or of each streamed delta), separate from content. Its tokens count as output. If max_tokens is small, a model can spend the whole budget on reasoning and return an empty content; give reasoning models enough room.

Image tokens#

On models with image input, each image counts toward prompt_tokens. A small image costs only a few tokens; from about 1024 pixels per side the count levels off, and one image is bounded at about 578 prompt tokens however many pixels it has. The hold for a request with images uses that bound per image. The Vision input guide covers formats and size limits.

Cancelled and interrupted requests#

If you close the connection while a response is streaming, the model stops and the request is billed for what was produced. When no final usage is available, we estimate it: the prompt at its upper bound, and the delivered text at about four characters per token. The estimate is never higher than the hold placed for the request. A request that fails before any content reaches you is not billed. A response that ends with no content and no usage counts as a failed request: you get an upstream_error and it is not billed.

Check usage over time#

The Usage page in the dashboard lists your requests with their token counts and cost.