Chat completions
Reference for POST /v1/chat/completions: every request parameter, the response object, the streaming chunk format and errors.
Creates a model response for a conversation. The request and response follow the OpenAI Chat Completions format.
POST https://usemodellane.com/v1/chat/completions
Send the conversation in messages and the model id in model; every other field is optional. The request and response follow the OpenAI Chat Completions format, so the official SDKs work with only the base URL and key changed.
With stream: true the reply arrives as server-sent events; the chunk format is under Stream chunk below, and the Streaming responses guide covers the details. Images are accepted only by models with image input, and only as base64 data URLs; see Vision input.
Request
modelstring requiredThe id of the model to use. See Models & Pricing for the list.
messagesarray requiredThe conversation so far, oldest message first. The API is stateless: send the full history on every request.
System message
rolestring requiredThe role of the message author.
contentstring requiredInstructions that set the assistant's behavior for the conversation.
namestringAn optional name for the participant.
User message
rolestring requiredThe role of the message author.
contentstring | array requiredThe user's message. Send a string, or an array of parts (`text` and `image_url`) for models that accept image input. Images are sent inline as base64 data URLs.
namestringAn optional name for the participant.
Assistant message
rolestring requiredThe role of the message author.
contentstring | nullA previous reply from the assistant. May be null when the turn only holds tool calls.
tool_callsarrayTool calls the assistant made in that turn, returned unchanged from the earlier response.
Tool message
rolestring requiredThe role of the message author.
contentstring requiredThe result of the tool call, usually a JSON string.
tool_call_idstring requiredThe id of the tool call this message answers.
streambooleanWhen true, the response is sent as server-sent events: a series of `data:` chunks that ends with `data: [DONE]`.
stream_optionsobjectOptions for a streamed response. Only valid when `stream` is true.
include_usagebooleanWhen true, a final chunk before `data: [DONE]` carries the token usage of the whole request.
max_tokensintegerThe maximum number of tokens to generate. When omitted, the model's default output limit applies.
max_completion_tokensintegerSame as `max_tokens`, under the newer name. Send one of the two.
temperaturenumberSampling temperature between 0 and 2. Higher values give more varied output, lower values more focused output.
top_pnumberNucleus sampling: only tokens within the top `top_p` probability mass are considered. Adjust this or `temperature`, not both.
stopstring | arrayUp to four sequences where the model stops generating. The stop sequence is not included in the output.
toolsarrayFunctions the model may call. Each entry has `type: "function"` and a `function` object.
typestring requiredThe tool type.
functionobject requiredThe function definition.
namestring requiredThe function name the model uses to call it.
descriptionstringWhat the function does; the model reads it to decide when to call it.
parametersobjectThe function arguments as a JSON Schema object.
tool_choicestring | objectControls tool use: `none` never calls a tool, `auto` lets the model decide, `required` forces a call. An object `{"type": "function", "function": {"name": …}}` forces one specific function.
response_formatobjectSet `{"type": "json_object"}` to make the model reply with valid JSON. Also ask for JSON in the prompt.
typestring requiredThe output format.
userstringA stable identifier for your end user, so you can tell requests from different users apart.
Examples
curl
curl https://usemodellane.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MODELLANE_API_KEY" \
-d '{
"model": "lane-1",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"stream": false
}'Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MODELLANE_API_KEY"],
base_url="https://usemodellane.com/v1",
)
response = client.chat.completions.create(
model="lane-1",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
stream=False,
)
print(response.choices[0].message.content)Node.js
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.MODELLANE_API_KEY,
baseURL: "https://usemodellane.com/v1",
})
const response = await client.chat.completions.create({
model: "lane-1",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Hello!" },
],
stream: false,
})
console.log(response.choices[0].message.content)Response
idstring requiredA unique id for this completion.
objectstring requiredThe object type.
createdinteger requiredUnix timestamp (seconds) of when the completion was created.
modelstring requiredThe model that produced the completion.
choicesarray requiredThe generated replies.
indexinteger requiredPosition of the choice in the list.
messageobject requiredThe assistant message.
rolestring requiredAlways `assistant`.
contentstring | null requiredThe reply text. Null when the turn only holds tool calls, or when the whole output budget went to reasoning.
reasoning_contentstringThe model's reasoning before the answer, on models that reason. Billed as output tokens. You do not need to send it back in later requests.
tool_callsarrayTool calls the model wants you to run.
idstring requiredThe id of the call. Send it back as `tool_call_id` with the result.
typestring requiredThe tool type.
functionobject requiredThe function the model wants to call.
namestring requiredThe function name.
argumentsstring requiredThe arguments as a JSON string. Validate them before you run the function.
finish_reasonstring | null requiredWhy generation stopped: a natural end or stop sequence, the token limit, or a tool call. Null in stream chunks until the last one.
usageobject requiredToken counts for the request. Billing uses these numbers.
prompt_tokensinteger requiredTokens in the prompt, including cached tokens.
completion_tokensinteger requiredTokens generated, including reasoning tokens.
total_tokensinteger requiredPrompt plus completion tokens.
prompt_tokens_detailsobjectBreakdown of prompt tokens.
cached_tokensintegerPrompt tokens read from the context cache, billed at the cached input price.
completion_tokens_detailsobjectBreakdown of completion tokens.
reasoning_tokensintegerCompletion tokens spent on reasoning. They are part of `completion_tokens`.
Stream chunk
With stream: true the response is a series of data: events, each holding one chunk object, and ends with data: [DONE].
idstring requiredThe completion id; the same in every chunk of one response.
objectstring requiredThe object type.
createdinteger requiredUnix timestamp (seconds) of when the completion was created.
modelstring requiredThe model that produced the completion.
choicesarray requiredThe new part of each reply. Empty in the final usage chunk.
indexinteger requiredPosition of the choice in the list.
deltaobject requiredThe text added since the previous chunk. Join the deltas to build the full message.
rolestringSent in the first chunk only.
contentstring | nullThe next piece of the reply text.
reasoning_contentstringThe model's reasoning before the answer, on models that reason. Billed as output tokens. You do not need to send it back in later requests.
tool_callsarrayPieces of tool calls. The first piece of a call carries `id` and the function name; later pieces with the same `index` append to `arguments`.
indexinteger requiredWhich call this piece belongs to.
idstring requiredThe id of the call. Send it back as `tool_call_id` with the result.
typestring requiredThe tool type.
functionobject requiredThe function the model wants to call.
namestring requiredThe function name.
argumentsstring requiredThe arguments as a JSON string. Validate them before you run the function.
finish_reasonstring | null requiredWhy generation stopped: a natural end or stop sequence, the token limit, or a tool call. Null in stream chunks until the last one.
usageobjectToken counts for the whole request. Only in the final chunk, and only when `stream_options.include_usage` is true.
prompt_tokensinteger requiredTokens in the prompt, including cached tokens.
completion_tokensinteger requiredTokens generated, including reasoning tokens.
total_tokensinteger requiredPrompt plus completion tokens.
prompt_tokens_detailsobjectBreakdown of prompt tokens.
cached_tokensintegerPrompt tokens read from the context cache, billed at the cached input price.
completion_tokens_detailsobjectBreakdown of completion tokens.
reasoning_tokensintegerCompletion tokens spent on reasoning. They are part of `completion_tokens`.
200
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1767225600,
"model": "lane-1",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help you today?" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 19,
"completion_tokens": 10,
"total_tokens": 29,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 }
}
}200 (stream)
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"role":"assistant","content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[],"usage":{"prompt_tokens":19,"completion_tokens":2,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}}}
data: [DONE]400
{
"error": {
"message": "Image URLs are not fetched. Send the image inline as a base64 data URL.",
"type": "invalid_request_error",
"code": "image_url_not_supported"
}
}Errors#
Errors return a JSON body of the form {"error": {"message", "type", "code"}}. The codes you are most likely to see on this endpoint:
400 invalid_request: a field has the wrong type or an invalid value. The message names the field.400 image_url_not_supported: an image was sent as a web link instead of a data URL.400 unsupported_content: the model does not accept images.400 context_length_exceeded: the prompt plusmax_tokensdoes not fit in the model's context length.402 insufficient_balance: your balance cannot cover the request's reservation.404 model_not_found: no model has this id. List ids with List models.429 rate_limit_exceeded,429 concurrency_limit,429 model_capacity: slow down and retry; see Rate limits.502 upstream_error,503 model_unavailable,504 upstream_timeout: the model could not answer. Wait a moment and retry.
Every code is listed on Error codes.