Messages
Reference for POST /v1/messages and count_tokens: request fields, content blocks, stream events and the Anthropic error shape.
Creates a model reply in the Anthropic Messages format, for the Anthropic SDKs and tools built on them.
POST https://usemodellane.com/v1/messages
Send the conversation in messages, the model id in model and the output limit in max_tokens. The request and response follow the Anthropic Messages format, so the official Anthropic SDKs and tools such as Claude Code work with the base URL https://usemodellane.com and your key. Models, prices, limits and billing are the same as for Chat completions.
Claude model names are not mapped to other models; use an id from Models & pricing. Fields not listed below are ignored. The Anthropic Messages API guide explains thinking, tool use and caching with examples.
Request
modelstring requiredA model id from Models & Pricing. Claude model names are not mapped to other models and return `not_found_error`.
max_tokensinteger requiredThe maximum number of tokens to generate, reasoning included.
messagesarray requiredThe conversation so far, alternating user and assistant turns.
User message
rolestring requiredThe role of the message author.
contentstring | array requiredA string, or an array of content blocks: `text`; `image` with `source.type` `base64` (`media_type` image/png, image/jpeg, image/webp or image/gif); and `tool_result` (`tool_use_id`, `content`, optional `is_error`). An image with `source.type` `url` is rejected with `image_url_not_supported`; `document` blocks and file ids are rejected.
Assistant message
rolestring requiredThe role of the message author.
contentstring | array requiredAn earlier reply: `text` and `tool_use` blocks (`id`, `name`, `input`). `thinking` and `redacted_thinking` blocks are accepted and dropped.
systemstring | arrayThe system prompt, as a string or an array of `text` blocks.
toolsarrayTools the model may call. Server tools (web search, code execution, bash and others with a `type`) are ignored.
namestring requiredThe tool name.
descriptionstringWhat the tool does and when to use it.
input_schemaobject requiredThe tool input as a JSON Schema object.
tool_choiceobject`type` is `auto`, `any`, `tool` (with `name`) or `none`. `disable_parallel_tool_use` is ignored.
temperaturenumberSampling temperature between 0 and 1.
top_pnumberNucleus sampling between 0 and 1.
stop_sequencesarrayUp to 4 strings that end the reply.
streambooleanSend the reply as server-sent events.
thinkingobjectAny `type` other than `disabled` returns the model's reasoning as `thinking` blocks. `budget_tokens` is not applied: models that reason do so on their own, and reasoning is billed as output tokens either way.
top_kintegerAccepted and ignored. Not sent to the model.
metadataobjectAccepted and ignored. Not stored.
cache_controlobjectAccepted and ignored. Caching is automatic, on any block.
Examples
curl
curl https://usemodellane.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $MODELLANE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "lane-1",
"max_tokens": 1024,
"system": "You are a helpful assistant.",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'Python
import os
import anthropic
client = anthropic.Anthropic(
api_key=os.environ["MODELLANE_API_KEY"],
base_url="https://usemodellane.com",
)
message = client.messages.create(
model="lane-1",
max_tokens=1024,
system="You are a helpful assistant.",
messages=[{"role": "user", "content": "Hello!"}],
)
print(message.content[0].text)Node.js
import Anthropic from "@anthropic-ai/sdk"
const client = new Anthropic({
apiKey: process.env.MODELLANE_API_KEY,
baseURL: "https://usemodellane.com",
})
const message = await client.messages.create({
model: "lane-1",
max_tokens: 1024,
system: "You are a helpful assistant.",
messages: [{ role: "user", content: "Hello!" }],
})
console.log(message.content[0].text)Response
idstring requiredA unique id for the message, starting with `msg_`.
typestring requiredThe object type.
rolestring requiredAlways `assistant`.
modelstring requiredThe model that produced the reply.
contentarray requiredContent blocks in order: thinking, text, then tool use.
Text
typestring requiredA text block.
textstring requiredThe reply text.
Thinking
typestring requiredReturned only when the request turned `thinking` on and the model produced reasoning.
thinkingstring requiredThe reasoning text.
signaturestring requiredAlways an empty string.
Tool use
typestring requiredA tool call.
idstring requiredSend it back as `tool_use_id` in a `tool_result` block.
namestring requiredThe tool the model wants to call.
inputobject requiredThe tool arguments.
stop_reasonstring requiredWhy the reply ended. A matched stop sequence reports `end_turn`.
stop_sequencenull requiredAlways null.
usageobject requiredToken counts for the request, the same numbers you are billed for.
input_tokensinteger requiredPrompt tokens not served from cache.
cache_read_input_tokensinteger requiredPrompt tokens served from cache.
cache_creation_input_tokensinteger requiredAlways 0: cache writes are not billed separately.
output_tokensinteger requiredTokens generated, reasoning included.
Stream events
With stream: true the reply arrives as named server-sent events in the Anthropic format, in this order:
message_starteventFirst event, with the message shell. Its usage counts are zero; the real counts arrive in `message_delta`.
content_block_starteventA content block starts: `text`, `thinking` or `tool_use`.
content_block_deltaeventA piece of the block: `text_delta`, `thinking_delta` or `input_json_delta` (tool arguments as partial JSON).
content_block_stopeventThe block is complete.
message_deltaeventThe `stop_reason` and the final usage.
message_stopeventLast event of a complete reply.
pingeventKeep-alive; ignore it.
erroreventThe request failed after the stream started. The stream ends without `message_stop`.
Count tokens
POST /v1/messages/count_tokens takes the same body without max_tokens and returns an estimate of the prompt size without running the model. It is free and does not touch your balance. The number is estimated from the size of the prompt, so it can differ from the prompt tokens you are billed for.
input_tokensinteger requiredThe estimated number of prompt tokens.
200
{
"id": "msg_123",
"type": "message",
"role": "assistant",
"model": "lane-1",
"content": [
{ "type": "text", "text": "Hello! How can I help you today?" }
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 19,
"cache_read_input_tokens": 0,
"cache_creation_input_tokens": 0,
"output_tokens": 10
}
}200 (stream)
event: message_start
data: {"type":"message_start","message":{"id":"msg_123","type":"message","role":"assistant","model":"lane-1","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"output_tokens":0}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello!"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":19,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":2}}
event: message_stop
data: {"type":"message_stop"}404
{
"type": "error",
"error": {
"type": "not_found_error",
"message": "The model 'claude-example' does not exist."
}
}Errors#
Errors use the Anthropic shape, {"type": "error", "error": {"type", "message"}}. The HTTP status is the same as on the other endpoints:
400 invalid_request_error: a missingmax_tokens, an invalid field, an image link instead of base64 data, an image sent to a model without image input, or a prompt that does not fit in the context length.401 authentication_error: the key is missing or wrong.402 billing_error: your balance cannot cover the request's reservation.404 not_found_error: the model id is not in the catalog, which includes every Claude model name.413 request_too_large: the body is over 8 MB.429 rate_limit_error: slow down and retry afterRetry-After; see Rate limits.500 api_error,502 api_error,503 overloaded_error,504 timeout_error: the model could not answer. Wait a moment and retry.
When streaming, an error after the first event arrives as an error event and the stream ends without message_stop.