Messages

Reference for POST /v1/messages and count_tokens: request fields, content blocks, stream events and the Anthropic error shape.

Creates a model reply in the Anthropic Messages format, for the Anthropic SDKs and tools built on them.

POST https://usemodellane.com/v1/messages

Send the conversation in messages, the model id in model and the output limit in max_tokens. The request and response follow the Anthropic Messages format, so the official Anthropic SDKs and tools such as Claude Code work with the base URL https://usemodellane.com and your key. Models, prices, limits and billing are the same as for Chat completions.

Claude model names are not mapped to other models; use an id from Models & pricing. Fields not listed below are ignored. The Anthropic Messages API guide explains thinking, tool use and caching with examples.

Request

  • model string required

    A model id from Models & Pricing. Claude model names are not mapped to other models and return `not_found_error`.

  • max_tokens integer required

    The maximum number of tokens to generate, reasoning included.

  • messages array required

    The conversation so far, alternating user and assistant turns.

    User message

    • role string required

      Possible values: user

      The role of the message author.

    • content string | array required

      A string, or an array of content blocks: `text`; `image` with `source.type` `base64` (`media_type` image/png, image/jpeg, image/webp or image/gif); and `tool_result` (`tool_use_id`, `content`, optional `is_error`). An image with `source.type` `url` is rejected with `image_url_not_supported`; `document` blocks and file ids are rejected.

    Assistant message

    • role string required

      Possible values: assistant

      The role of the message author.

    • content string | array required

      An earlier reply: `text` and `tool_use` blocks (`id`, `name`, `input`). `thinking` and `redacted_thinking` blocks are accepted and dropped.

  • system string | array

    The system prompt, as a string or an array of `text` blocks.

  • tools array

    Tools the model may call. Server tools (web search, code execution, bash and others with a `type`) are ignored.

    • name string required

      The tool name.

    • description string

      What the tool does and when to use it.

    • input_schema object required

      The tool input as a JSON Schema object.

  • tool_choice object

    `type` is `auto`, `any`, `tool` (with `name`) or `none`. `disable_parallel_tool_use` is ignored.

  • temperature number

    Sampling temperature between 0 and 1.

  • top_p number

    Nucleus sampling between 0 and 1.

  • stop_sequences array

    Up to 4 strings that end the reply.

  • stream boolean

    Default: false

    Send the reply as server-sent events.

  • thinking object

    Any `type` other than `disabled` returns the model's reasoning as `thinking` blocks. `budget_tokens` is not applied: models that reason do so on their own, and reasoning is billed as output tokens either way.

  • top_k integer

    Accepted and ignored. Not sent to the model.

  • metadata object

    Accepted and ignored. Not stored.

  • cache_control object

    Accepted and ignored. Caching is automatic, on any block.

Examples

curl

curl https://usemodellane.com/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: $MODELLANE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "lane-1",
    "max_tokens": 1024,
    "system": "You are a helpful assistant.",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Python

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com",
)

message = client.messages.create(
    model="lane-1",
    max_tokens=1024,
    system="You are a helpful assistant.",
    messages=[{"role": "user", "content": "Hello!"}],
)

print(message.content[0].text)

Node.js

import Anthropic from "@anthropic-ai/sdk"

const client = new Anthropic({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com",
})

const message = await client.messages.create({
  model: "lane-1",
  max_tokens: 1024,
  system: "You are a helpful assistant.",
  messages: [{ role: "user", content: "Hello!" }],
})

console.log(message.content[0].text)

Response

  • id string required

    A unique id for the message, starting with `msg_`.

  • type string required

    Possible values: message

    The object type.

  • role string required

    Possible values: assistant

    Always `assistant`.

  • model string required

    The model that produced the reply.

  • content array required

    Content blocks in order: thinking, text, then tool use.

    Text

    • type string required

      Possible values: text

      A text block.

    • text string required

      The reply text.

    Thinking

    • type string required

      Possible values: thinking

      Returned only when the request turned `thinking` on and the model produced reasoning.

    • thinking string required

      The reasoning text.

    • signature string required

      Always an empty string.

    Tool use

    • type string required

      Possible values: tool_use

      A tool call.

    • id string required

      Send it back as `tool_use_id` in a `tool_result` block.

    • name string required

      The tool the model wants to call.

    • input object required

      The tool arguments.

  • stop_reason string required

    Possible values: end_turn, max_tokens, tool_use, refusal

    Why the reply ended. A matched stop sequence reports `end_turn`.

  • stop_sequence null required

    Always null.

  • usage object required

    Token counts for the request, the same numbers you are billed for.

    • input_tokens integer required

      Prompt tokens not served from cache.

    • cache_read_input_tokens integer required

      Prompt tokens served from cache.

    • cache_creation_input_tokens integer required

      Always 0: cache writes are not billed separately.

    • output_tokens integer required

      Tokens generated, reasoning included.

Stream events

With stream: true the reply arrives as named server-sent events in the Anthropic format, in this order:

  • message_start event

    First event, with the message shell. Its usage counts are zero; the real counts arrive in `message_delta`.

  • content_block_start event

    A content block starts: `text`, `thinking` or `tool_use`.

  • content_block_delta event

    A piece of the block: `text_delta`, `thinking_delta` or `input_json_delta` (tool arguments as partial JSON).

  • content_block_stop event

    The block is complete.

  • message_delta event

    The `stop_reason` and the final usage.

  • message_stop event

    Last event of a complete reply.

  • ping event

    Keep-alive; ignore it.

  • error event

    The request failed after the stream started. The stream ends without `message_stop`.

Count tokens

POST /v1/messages/count_tokens takes the same body without max_tokens and returns an estimate of the prompt size without running the model. It is free and does not touch your balance. The number is estimated from the size of the prompt, so it can differ from the prompt tokens you are billed for.

  • input_tokens integer required

    The estimated number of prompt tokens.

200

{
  "id": "msg_123",
  "type": "message",
  "role": "assistant",
  "model": "lane-1",
  "content": [
    { "type": "text", "text": "Hello! How can I help you today?" }
  ],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 19,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0,
    "output_tokens": 10
  }
}

200 (stream)

event: message_start
data: {"type":"message_start","message":{"id":"msg_123","type":"message","role":"assistant","model":"lane-1","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello!"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":19,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":2}}

event: message_stop
data: {"type":"message_stop"}

404

{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "The model 'claude-example' does not exist."
  }
}

Errors#

Errors use the Anthropic shape, {"type": "error", "error": {"type", "message"}}. The HTTP status is the same as on the other endpoints:

  • 400 invalid_request_error: a missing max_tokens, an invalid field, an image link instead of base64 data, an image sent to a model without image input, or a prompt that does not fit in the context length.
  • 401 authentication_error: the key is missing or wrong.
  • 402 billing_error: your balance cannot cover the request's reservation.
  • 404 not_found_error: the model id is not in the catalog, which includes every Claude model name.
  • 413 request_too_large: the body is over 8 MB.
  • 429 rate_limit_error: slow down and retry after Retry-After; see Rate limits.
  • 500 api_error, 502 api_error, 503 overloaded_error, 504 timeout_error: the model could not answer. Wait a moment and retry.

When streaming, an error after the first event arrives as an error event and the stream ends without message_stop.