Models & pricing
Every Modellane model with its context length, output limit, supported features and per-token price, read live from the model catalog.
This page lists the models you can call, what each one supports and what it costs per million tokens. The table is generated from the live catalog, so it always matches what the API bills.
The tables on this page are generated from the live model catalog, the same catalog the API bills from. There are 3 models available right now. Use the value in the Model column as the model parameter of a request.
Models#
| Model | Name | Context length | Max output | Features |
|---|---|---|---|---|
deepseek-roleplay | DeepSeek-Roleplay | 1,048,576 | 393,216 | Streaming, Tool calls, JSON output, Reasoning |
lane-1 | Lane 1 | 262,144 | 262,134 | Streaming, Tool calls, JSON output, Vision, Reasoning |
lane-1-pro | Lane 1 Pro | 1,000,000 | 999,990 | Streaming, Tool calls, JSON output, Reasoning |
- Context length is the total number of tokens a request can use: the prompt plus the generated output.
- Max output is the largest value you can pass as
max_tokens. Larger values are lowered to this limit. - Features lists what the model supports, such as streaming, tool calls, JSON output and image input. A request that uses an unsupported input type fails with
unsupported_content.
You can also list the models your key can call with GET /v1/models.
Pricing#
| Model | Input (per 1M tokens) | Cached input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|---|
deepseek-roleplay | $0.50 | $0.10 | $1.00 |
lane-1 | $1.50 | $0.15 | $4.50 |
lane-1-pro | $4.50 | $0.45 | $7.50 |
Prices are in US dollars per one million tokens. A request costs:
text
cost = (input tokens - cached tokens) x input price
+ cached tokens x cached input price
+ output tokens x output price- Cached input applies to the part of the prompt that was served from cache, reported as
cached_tokensin the usage block. Caching is automatic; see Context caching. - Reasoning tokens are output tokens. When a model thinks before it answers, those tokens are part of
completion_tokensand are billed at the output price. - A dash means the model has no separate price for that row.
The exact token counts of every request are in its usage block. See Tokens & token usage.
Deduction rules#
The API checks your balance before it forwards a request:
- Hold. We place a temporary hold for the most the request could cost: up to one token per byte of prompt text (plus a fixed allowance per image) at the uncached input price, plus the full output budget at the output price. The hold is an upper bound, not an estimate of the charge.
- Charge. When the response is finished, we charge the actual cost from the reported usage and release the rest of the hold. The charge is never higher than the hold.
- No content, no charge. If the request fails before any content reaches you, the hold is released in full.
If you omit max_tokens, the output budget is the model's default output limit, which is lower than its Max output. Setting max_tokens close to what you really need keeps the hold small, which matters when many requests run in parallel on a small balance.
Prices can change. A price change applies to requests sent after it takes effect and is listed in the Changelog.