Chat completions

Reference for POST /v1/chat/completions: every request parameter, the response object, the streaming chunk format and errors.

Creates a model response for a conversation. The request and response follow the OpenAI Chat Completions format.

POST https://usemodellane.com/v1/chat/completions

Send the conversation in messages and the model id in model; every other field is optional. The request and response follow the OpenAI Chat Completions format, so the official SDKs work with only the base URL and key changed.

With stream: true the reply arrives as server-sent events; the chunk format is under Stream chunk below, and the Streaming responses guide covers the details. Images are accepted only by models with image input, and only as base64 data URLs; see Vision input.

Request

  • model string required

    The id of the model to use. See Models & Pricing for the list.

  • messages array required

    The conversation so far, oldest message first. The API is stateless: send the full history on every request.

    System message

    • role string required

      Possible values: system

      The role of the message author.

    • content string required

      Instructions that set the assistant's behavior for the conversation.

    • name string

      An optional name for the participant.

    User message

    • role string required

      Possible values: user

      The role of the message author.

    • content string | array required

      The user's message. Send a string, or an array of parts (`text` and `image_url`) for models that accept image input. Images are sent inline as base64 data URLs.

    • name string

      An optional name for the participant.

    Assistant message

    • role string required

      Possible values: assistant

      The role of the message author.

    • content string | null

      A previous reply from the assistant. May be null when the turn only holds tool calls.

    • tool_calls array

      Tool calls the assistant made in that turn, returned unchanged from the earlier response.

    Tool message

    • role string required

      Possible values: tool

      The role of the message author.

    • content string required

      The result of the tool call, usually a JSON string.

    • tool_call_id string required

      The id of the tool call this message answers.

  • stream boolean

    Default: false

    When true, the response is sent as server-sent events: a series of `data:` chunks that ends with `data: [DONE]`.

  • stream_options object

    Options for a streamed response. Only valid when `stream` is true.

    • include_usage boolean

      When true, a final chunk before `data: [DONE]` carries the token usage of the whole request.

  • max_tokens integer

    The maximum number of tokens to generate. When omitted, the model's default output limit applies.

  • max_completion_tokens integer

    Same as `max_tokens`, under the newer name. Send one of the two.

  • temperature number

    Sampling temperature between 0 and 2. Higher values give more varied output, lower values more focused output.

  • top_p number

    Nucleus sampling: only tokens within the top `top_p` probability mass are considered. Adjust this or `temperature`, not both.

  • stop string | array

    Up to four sequences where the model stops generating. The stop sequence is not included in the output.

  • tools array

    Functions the model may call. Each entry has `type: "function"` and a `function` object.

    • type string required

      Possible values: function

      The tool type.

    • function object required

      The function definition.

      • name string required

        The function name the model uses to call it.

      • description string

        What the function does; the model reads it to decide when to call it.

      • parameters object

        The function arguments as a JSON Schema object.

  • tool_choice string | object

    Possible values: none, auto, required

    Controls tool use: `none` never calls a tool, `auto` lets the model decide, `required` forces a call. An object `{"type": "function", "function": {"name": …}}` forces one specific function.

  • response_format object

    Set `{"type": "json_object"}` to make the model reply with valid JSON. Also ask for JSON in the prompt.

    • type string required

      Possible values: text, json_object

      The output format.

  • user string

    A stable identifier for your end user, so you can tell requests from different users apart.

Examples

curl

curl https://usemodellane.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MODELLANE_API_KEY" \
  -d '{
    "model": "lane-1",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "stream": false
  }'

Python

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
)

response = client.chat.completions.create(
    model="lane-1",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
    stream=False,
)

print(response.choices[0].message.content)

Node.js

import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
})

const response = await client.chat.completions.create({
  model: "lane-1",
  messages: [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Hello!" },
  ],
  stream: false,
})

console.log(response.choices[0].message.content)

Response

  • id string required

    A unique id for this completion.

  • object string required

    Possible values: chat.completion

    The object type.

  • created integer required

    Unix timestamp (seconds) of when the completion was created.

  • model string required

    The model that produced the completion.

  • choices array required

    The generated replies.

    • index integer required

      Position of the choice in the list.

    • message object required

      The assistant message.

      • role string required

        Possible values: assistant

        Always `assistant`.

      • content string | null required

        The reply text. Null when the turn only holds tool calls, or when the whole output budget went to reasoning.

      • reasoning_content string

        The model's reasoning before the answer, on models that reason. Billed as output tokens. You do not need to send it back in later requests.

      • tool_calls array

        Tool calls the model wants you to run.

        • id string required

          The id of the call. Send it back as `tool_call_id` with the result.

        • type string required

          Possible values: function

          The tool type.

        • function object required

          The function the model wants to call.

          • name string required

            The function name.

          • arguments string required

            The arguments as a JSON string. Validate them before you run the function.

    • finish_reason string | null required

      Possible values: stop, length, tool_calls

      Why generation stopped: a natural end or stop sequence, the token limit, or a tool call. Null in stream chunks until the last one.

  • usage object required

    Token counts for the request. Billing uses these numbers.

    • prompt_tokens integer required

      Tokens in the prompt, including cached tokens.

    • completion_tokens integer required

      Tokens generated, including reasoning tokens.

    • total_tokens integer required

      Prompt plus completion tokens.

    • prompt_tokens_details object

      Breakdown of prompt tokens.

      • cached_tokens integer

        Prompt tokens read from the context cache, billed at the cached input price.

    • completion_tokens_details object

      Breakdown of completion tokens.

      • reasoning_tokens integer

        Completion tokens spent on reasoning. They are part of `completion_tokens`.

Stream chunk

With stream: true the response is a series of data: events, each holding one chunk object, and ends with data: [DONE].

  • id string required

    The completion id; the same in every chunk of one response.

  • object string required

    Possible values: chat.completion.chunk

    The object type.

  • created integer required

    Unix timestamp (seconds) of when the completion was created.

  • model string required

    The model that produced the completion.

  • choices array required

    The new part of each reply. Empty in the final usage chunk.

    • index integer required

      Position of the choice in the list.

    • delta object required

      The text added since the previous chunk. Join the deltas to build the full message.

      • role string

        Possible values: assistant

        Sent in the first chunk only.

      • content string | null

        The next piece of the reply text.

      • reasoning_content string

        The model's reasoning before the answer, on models that reason. Billed as output tokens. You do not need to send it back in later requests.

      • tool_calls array

        Pieces of tool calls. The first piece of a call carries `id` and the function name; later pieces with the same `index` append to `arguments`.

        • index integer required

          Which call this piece belongs to.

        • id string required

          The id of the call. Send it back as `tool_call_id` with the result.

        • type string required

          Possible values: function

          The tool type.

        • function object required

          The function the model wants to call.

          • name string required

            The function name.

          • arguments string required

            The arguments as a JSON string. Validate them before you run the function.

    • finish_reason string | null required

      Possible values: stop, length, tool_calls

      Why generation stopped: a natural end or stop sequence, the token limit, or a tool call. Null in stream chunks until the last one.

  • usage object

    Token counts for the whole request. Only in the final chunk, and only when `stream_options.include_usage` is true.

    • prompt_tokens integer required

      Tokens in the prompt, including cached tokens.

    • completion_tokens integer required

      Tokens generated, including reasoning tokens.

    • total_tokens integer required

      Prompt plus completion tokens.

    • prompt_tokens_details object

      Breakdown of prompt tokens.

      • cached_tokens integer

        Prompt tokens read from the context cache, billed at the cached input price.

    • completion_tokens_details object

      Breakdown of completion tokens.

      • reasoning_tokens integer

        Completion tokens spent on reasoning. They are part of `completion_tokens`.

200

{
  "id": "chatcmpl-123",
  "object": "chat.completion",
  "created": 1767225600,
  "model": "lane-1",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello! How can I help you today?" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 10,
    "total_tokens": 29,
    "prompt_tokens_details": { "cached_tokens": 0 },
    "completion_tokens_details": { "reasoning_tokens": 0 }
  }
}

200 (stream)

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"role":"assistant","content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[],"usage":{"prompt_tokens":19,"completion_tokens":2,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}}}

data: [DONE]

400

{
  "error": {
    "message": "Image URLs are not fetched. Send the image inline as a base64 data URL.",
    "type": "invalid_request_error",
    "code": "image_url_not_supported"
  }
}

Errors#

Errors return a JSON body of the form {"error": {"message", "type", "code"}}. The codes you are most likely to see on this endpoint:

  • 400 invalid_request: a field has the wrong type or an invalid value. The message names the field.
  • 400 image_url_not_supported: an image was sent as a web link instead of a data URL.
  • 400 unsupported_content: the model does not accept images.
  • 400 context_length_exceeded: the prompt plus max_tokens does not fit in the model's context length.
  • 402 insufficient_balance: your balance cannot cover the request's reservation.
  • 404 model_not_found: no model has this id. List ids with List models.
  • 429 rate_limit_exceeded, 429 concurrency_limit, 429 model_capacity: slow down and retry; see Rate limits.
  • 502 upstream_error, 503 model_unavailable, 504 upstream_timeout: the model could not answer. Wait a moment and retry.

Every code is listed on Error codes.