Responses

Reference for POST /v1/responses: request fields, the response object, streaming events and the fields that are not supported.

Creates a model response in the OpenAI Responses format. Nothing is stored, so every request carries the full conversation.

POST https://usemodellane.com/v1/responses

Send the prompt in input and the model id in model. The request and response follow the OpenAI Responses format, so client.responses.create in the official SDKs and tools such as Codex work with only the base URL and key changed. Models, prices, limits and billing are the same as for Chat completions.

Hosted tools that run on OpenAI's servers (web search, file search, code interpreter, image generation) are removed from the request without an error. Fields not listed below are ignored. The Responses API guide covers conversations, tools and streaming with examples.

Request

  • model string required

    The id of the model to use, the same id as in a chat completion request.

  • input string | array required

    The prompt as a string, or the whole conversation as an array of items. Responses are never stored, so send the full history with every request. `custom_tool_call` and `custom_tool_call_output` items are accepted as well; earlier `reasoning` items are accepted and dropped.

    Message

    • type string

      Possible values: message

      The item type. May be left out for messages.

    • role string required

      Possible values: user, assistant, system, developer

      The author of the message. `developer` is treated like `system`.

    • content string | array required

      A string, or an array of parts: `input_text` (`text`), `output_text` (`text`, for earlier assistant turns) and `input_image` (`image_url` as a base64 data URL, with an optional `detail`). Image links and `file_id` are rejected with `image_url_not_supported`; `input_file` and `input_audio` with `unsupported_content`.

    Function call

    • type string required

      Possible values: function_call

      A tool call the model made in an earlier turn, sent back unchanged.

    • call_id string required

      The id that links the call to its output.

    • name string required

      The function name.

    • arguments string required

      The arguments as a JSON string.

    Function call output

    • type string required

      Possible values: function_call_output

      The result of a tool call, produced by your code.

    • call_id string required

      The `call_id` of the call this output answers.

    • output string | array required

      The result as a string, or an array of `input_text` and `input_image` parts. Images need a model with image input.

  • instructions string

    A system message placed before the input.

  • tools array

    Tools the model may call. Function tools work as in Chat Completions. Custom tools receive one free-text input; a grammar in the tool is passed to the model as a hint and is not enforced. Tools inside a `namespace` are flattened to plain names. Hosted tools such as web search, file search and code interpreter are ignored.

    • type string required

      Possible values: function, custom, namespace

      The kind of tool.

    • name string required

      The tool name the model uses to call it.

    • description string

      What the tool does and when to use it.

    • parameters object

      The arguments of a function tool as a JSON Schema object.

  • tool_choice string | object

    Possible values: none, auto, required

    Default: auto

    Whether the model may, must or must not call tools. To force one tool, send `{"type": "function", "name": "..."}`.

  • max_output_tokens integer

    The maximum number of tokens to generate, reasoning included. Works like `max_tokens` in Chat Completions.

  • temperature number

    Sampling temperature, as in Chat Completions.

  • top_p number

    Nucleus sampling, as in Chat Completions.

  • text object

    Output format. `text.format.type` is `text`, `json_object` or `json_schema`, as `response_format` in Chat Completions. `text.verbosity` is ignored.

  • stream boolean

    Default: false

    Send the response as server-sent events. The final `response.completed` event carries the usage.

  • metadata object

    Up to 16 string pairs, returned unchanged in the response. Not stored.

  • store boolean

    Accepted, but responses are never stored; the response always reports `false`.

  • reasoning object

    Accepted and ignored. Reasoning effort cannot be set; models that reason do so on their own, and their reasoning comes back as a summary item.

  • parallel_tool_calls boolean

    Accepted and ignored. Echoed in the response.

  • include array

    Accepted and ignored. Encrypted reasoning is never returned.

  • prompt_cache_key string

    Accepted and ignored. Repeated prefixes are cached automatically.

  • previous_response_id string

    Not supported, because responses are not stored. Requests with `previous_response_id`, `conversation`, `prompt` or `background: true` fail with `unsupported_parameter`.

Examples

curl

curl https://usemodellane.com/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MODELLANE_API_KEY" \
  -d '{
    "model": "lane-1",
    "instructions": "You are a helpful assistant.",
    "input": "Hello!"
  }'

Python

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
)

response = client.responses.create(
    model="lane-1",
    instructions="You are a helpful assistant.",
    input="Hello!",
)

print(response.output_text)

Node.js

import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
})

const response = await client.responses.create({
  model: "lane-1",
  instructions: "You are a helpful assistant.",
  input: "Hello!",
})

console.log(response.output_text)

Response

  • id string required

    A unique id for the response, starting with `resp_`.

  • object string required

    Possible values: response

    The object type.

  • created_at integer required

    Unix timestamp (seconds) of when the response was created.

  • status string required

    Possible values: completed, incomplete

    `incomplete` when the output stopped early; `incomplete_details.reason` says why.

  • incomplete_details object | null required

    `reason` is `max_output_tokens` or `content_filter` when the status is `incomplete`.

  • model string required

    The model that produced the response.

  • output array required

    The output items in order: reasoning, the message, then tool calls.

    Message

    • type string required

      Possible values: message

      The item type.

    • id string required

      The item id.

    • role string required

      Possible values: assistant

      Always `assistant`.

    • content array required

      `output_text` parts with the reply in `text`.

    Reasoning

    • type string required

      Possible values: reasoning

      Returned when the model produced reasoning.

    • summary array required

      `summary_text` parts with the reasoning text.

    • encrypted_content null required

      Always null.

    Function call

    • type string required

      Possible values: function_call, custom_tool_call

      `custom_tool_call` when the tool was declared as a custom tool.

    • call_id string required

      Send it back in the matching `function_call_output` item.

    • name string required

      The tool the model wants to call.

    • arguments string required

      The arguments as a JSON string (`input` for custom tools).

  • usage object required

    Token counts for the request, the same numbers you are billed for.

    • input_tokens integer required

      Tokens in the prompt.

    • input_tokens_details.cached_tokens integer required

      Prompt tokens served from cache.

    • output_tokens integer required

      Tokens generated, reasoning included.

    • output_tokens_details.reasoning_tokens integer required

      Generated tokens spent on reasoning.

    • total_tokens integer required

      Input plus output tokens.

  • store boolean required

    Possible values: false

    Always false: nothing is stored.

  • metadata object required

    The metadata you sent.

Stream events

With stream: true every event is an event: line and a data: line holding a JSON object with type and an increasing sequence_number. The main events, in order:

  • response.created event

    First event. Carries the response object with status `in_progress`; `response.in_progress` follows.

  • response.output_item.added event

    A new output item starts: a message, a reasoning item or a function call (with empty arguments).

  • response.content_part.added event

    A text part starts inside a message item.

  • response.output_text.delta event

    A piece of the reply text in `delta`.

  • response.output_text.done event

    The full text of the part; followed by `response.content_part.done`.

  • response.reasoning_summary_text.delta event

    A piece of the reasoning text. Framed by `response.reasoning_summary_part.added` and `.done`.

  • response.function_call_arguments.delta event

    A piece of a function call's arguments; `.done` carries the complete arguments.

  • response.output_item.done event

    The finished item. Function calls are complete only in this event.

  • response.completed event

    Last event. Carries the full response with `output` and `usage`. `response.incomplete` replaces it when the output stopped early.

  • response.failed event

    The request failed after the stream started. `response.error.code` names the error.

200

{
  "id": "resp_123",
  "object": "response",
  "created_at": 1767225600,
  "status": "completed",
  "error": null,
  "incomplete_details": null,
  "model": "lane-1",
  "output": [
    {
      "id": "msg_123",
      "type": "message",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Hello! How can I help you today?", "annotations": [], "logprobs": [] }
      ]
    }
  ],
  "usage": {
    "input_tokens": 19,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 10,
    "output_tokens_details": { "reasoning_tokens": 0 },
    "total_tokens": 29
  },
  "store": false,
  "metadata": {}
}

200 (stream)

event: response.created
data: {"type":"response.created","sequence_number":0,"response":{"id":"resp_123","object":"response","status":"in_progress"}}

event: response.output_text.delta
data: {"type":"response.output_text.delta","sequence_number":4,"item_id":"msg_123","output_index":0,"content_index":0,"delta":"Hello"}

event: response.completed
data: {"type":"response.completed","sequence_number":9,"response":{"id":"resp_123","object":"response","status":"completed","usage":{"input_tokens":19,"output_tokens":2,"total_tokens":21}}}

400

{
  "error": {
    "message": "previous_response_id is not supported: responses are never stored.",
    "type": "invalid_request_error",
    "param": "previous_response_id",
    "code": "unsupported_parameter"
  }
}

Errors#

Errors return {"error": {"message", "type", "param", "code"}}, with param naming the Responses field. The codes are the same as for Chat Completions, plus:

  • 400 unsupported_parameter: the request uses stored state (previous_response_id, conversation, prompt, background) or an item_reference item.
  • 400 image_url_not_supported: an input_image holds a web link or a file_id instead of a base64 data URL.
  • 400 unsupported_content: an input_file or input_audio part, or an image sent to a model without image input.

When streaming, an error after the first event arrives as a response.failed event with response.error.code. Every code is listed on Error codes.