Anthropic Messages API

Call Modellane in the Anthropic Messages format: base URL, x-api-key, how fields map, thinking blocks, tool use and the limits.

Apps built on the Anthropic SDKs can call Modellane models through the Messages API. Change the base URL and the key, and use a Modellane model id.

Modellane serves the Anthropic Messages API at POST /v1/messages, next to the OpenAI formats. Apps built on the Anthropic SDKs, and tools such as Claude Code, can call Modellane models with two changes: the base URL and the key. Every model works in this format, at the same prices and limits as in Chat Completions.

Base URL and key#

ParamValue
Base URLhttps://usemodellane.com (no /v1: the SDKs add /v1/messages)
API keyA key from API keys, in x-api-key or as Authorization: Bearer
modellane-1, or any id from Models & pricing

The anthropic-version and anthropic-beta headers are accepted and not needed. Claude model names are not mapped to other models: a request for a Claude model gets a 404 not_found_error.

curl

curl https://usemodellane.com/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: $MODELLANE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "lane-1",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Python

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com",
)

message = client.messages.create(
    model="lane-1",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello!"}],
)
print(message.content[0].text)

Node.js

import Anthropic from "@anthropic-ai/sdk"

const client = new Anthropic({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com",
})

const message = await client.messages.create({
  model: "lane-1",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Hello!" }],
})
console.log(message.content[0].text)

How requests are handled#

A Messages request runs on the same engine as a chat completion. Billing, rate limits, balance holds and errors are the same; only the request and response shapes follow the Anthropic format.

Anthropic fieldBehavior
max_tokensRequired, as in the Anthropic API. Includes reasoning tokens.
systemA string or text blocks; becomes the system prompt.
text, image blocksSupported. Images only with source.type base64.
tool_use, tool_resultSupported, including is_error and images inside a tool result.
tools, tool_choiceYour own tools with input_schema; auto, any, tool and none.
temperature, top_pSupported, between 0 and 1.
stop_sequencesUp to 4 strings.
streamSupported, with the Anthropic event sequence.
thinkingChooses whether reasoning is returned; the budget is not applied.
top_k, metadata, cache_controlAccepted and ignored.

Thinking#

Models that reason do so on their own; the API cannot switch reasoning on or off or limit it. The thinking field only decides whether you see it: with any type other than disabled, the reasoning comes back as a thinking block before the text, with an empty signature. Without it, the reasoning is not returned. Either way, reasoning tokens are part of output_tokens and are billed as output.

Thinking blocks from earlier turns can be sent back in the history; they are dropped before the request reaches the model.

Tool use#

Tools work as in the Anthropic API. The model replies with tool_use blocks and stop_reason: "tool_use"; your code runs the tool and sends a tool_result block with the same tool_use_id in the next user message.

Python

response = client.messages.create(
    model="lane-1",
    max_tokens=1024,
    tools=[{
        "name": "get_weather",
        "description": "Return the current weather for a city.",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }],
    messages=[{"role": "user", "content": "What is the weather in Paris?"}],
)
for block in response.content:
    if block.type == "tool_use":
        print(block.name, block.input)

Pick a model that lists Tool calls in Models & pricing. Server tools that run on Anthropic's side, such as web search, code execution and the bash and text editor tools, are removed from the request without an error.

Usage and caching#

The usage block maps the billed tokens to Anthropic names: input_tokens are prompt tokens not served from cache, cache_read_input_tokens are cached prompt tokens, and output_tokens include reasoning. Caching is automatic, so cache_control is not needed and cache_creation_input_tokens is always 0. See Context caching.

In a stream, message_start carries zero usage; the real counts arrive in message_delta at the end.

Count tokens#

POST /v1/messages/count_tokens returns {"input_tokens": n} for a request body without running the model. It is free. The number is an estimate from the prompt size, so it can differ from the prompt tokens you are billed for.

Errors#

Errors use the Anthropic shape, {"type": "error", "error": {"type", "message"}}, with the same HTTP status as on the other endpoints:

StatusError typeCommon cause
400invalid_request_errorA missing or invalid field, an image link, a prompt longer than the context
401authentication_errorA missing or wrong key
402billing_errorThe balance cannot cover the request
404not_found_errorA model id that is not in the catalog
413request_too_largeA body over 8 MB
429rate_limit_errorToo many requests; wait for Retry-After
500, 502api_errorThe model could not answer
503overloaded_errorThe model is unavailable for a moment
504timeout_errorThe model took too long

The Anthropic SDKs retry 429 and 5xx errors on their own and do not retry 402.

Limits#

  • stop_sequence in the response is always null; a matched stop sequence reports stop_reason: "end_turn".
  • document and search_result blocks, Files API ids and image URLs are rejected.
  • The Message Batches API is not available.
  • The same request limits as Chat Completions apply: up to 2048 messages, 16 images of up to 5 MB each, 128 tools and an 8 MB body. Error messages can name fields in Chat Completions terms, such as messages[3].
  • Tool arguments the model returns as invalid JSON arrive as an empty input object.

The full field list is on Messages.