OpenAI Responses API

Use the OpenAI Responses format with Modellane: input items, function tools, streaming events, and what changes for Codex users.

The Responses API is the newer OpenAI request format, used by Codex and recent SDK features. Modellane serves it statelessly on the same models and prices.

Modellane serves the OpenAI Responses API at POST /v1/responses, on the same base URL, keys, models and prices as Chat Completions. It is the format Codex uses, and the one behind client.responses in the OpenAI SDKs. Responses are stateless: nothing is stored, so every request carries the whole conversation.

Make a request#

curl

curl https://usemodellane.com/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MODELLANE_API_KEY" \
  -d '{
    "model": "lane-1",
    "instructions": "You are a helpful assistant.",
    "input": "Hello!"
  }'

Python

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
)

response = client.responses.create(
    model="lane-1",
    instructions="You are a helpful assistant.",
    input="Hello!",
)
print(response.output_text)

Node.js

import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
})

const response = await client.responses.create({
  model: "lane-1",
  instructions: "You are a helpful assistant.",
  input: "Hello!",
})
console.log(response.output_text)

instructions becomes the system prompt. input is a string for a single prompt, or an array of items for a conversation.

Conversations without stored state#

Because nothing is stored, previous_response_id, conversation, prompt and background: true are rejected with unsupported_parameter, and there is no endpoint to fetch or delete an earlier response. Keep the history yourself and send it in input: user and assistant message items, plus the function_call and function_call_output items of earlier tool calls.

Python

history = [{"role": "user", "content": "Name a prime number."}]
first = client.responses.create(model="lane-1", input=history)

history += [
    {"role": "assistant", "content": first.output_text},
    {"role": "user", "content": "And the next one?"},
]
second = client.responses.create(model="lane-1", input=history)
print(second.output_text)

Function tools#

Declare tools with type: "function", a name and a JSON Schema in parameters. The model replies with function_call items; run the function and send the result back as a function_call_output item with the same call_id, together with the call itself.

Python

tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Return the current weather for a city.",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
}]

input_items = [{"role": "user", "content": "What is the weather in Paris?"}]
response = client.responses.create(model="lane-1", input=input_items, tools=tools)

for item in response.output:
    if item.type == "function_call":
        input_items.append(item)
        input_items.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": "Sunny, 22 C",
        })

final = client.responses.create(model="lane-1", input=input_items, tools=tools)
print(final.output_text)

Custom tools (type: "custom") receive one free-text input. Their grammar is passed to the model as a hint, not enforced. Tools grouped in a namespace are sent to the model under flat names and come back with their namespace. Hosted tools, such as web search, file search, code interpreter and image generation, are removed from the request without an error, because they run on OpenAI's servers.

Streaming#

With stream: true the response arrives as named server-sent events. Text comes in response.output_text.delta, tool arguments in response.function_call_arguments.delta, and the last event is response.completed with the full response and its usage. If the output stops early, response.incomplete replaces it; an error after the stream started arrives as response.failed. The SDKs handle the events for you:

Python

stream = client.responses.create(model="lane-1", input="Count to five.", stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)

The full event list is on Responses.

Reasoning#

Models that reason do so on their own; reasoning.effort and reasoning.summary have no effect. When the model produces reasoning, it comes back as a reasoning item whose summary holds the text, and the tokens count as output tokens. Encrypted reasoning is never returned, and reasoning items sent back in input are dropped.

Notes for Codex#

  • Set model_context_window in the Codex config. Codex does not read context lengths from a custom provider and otherwise assumes a fixed size; see Codex.
  • Codex sends its web search tool by default. It is skipped, so the model answers without web access.
  • Codex already sends the full history with store: false, which is what this endpoint needs.
  • Images from Codex's image viewing tool need a model with Vision; on a text-only model the request fails with unsupported_content.
  • Up to 128 tools fit in one request. A setup with many MCP servers can go over that and get a 400 error.

Other fields#

  • max_output_tokens works like max_tokens in Chat Completions; when it is reached, the status is incomplete with reason max_output_tokens.
  • text.format takes text, json_object or json_schema, like response_format. See JSON output.
  • metadata is returned unchanged; store is always reported as false.
  • include, parallel_tool_calls, prompt_cache_key, service_tier, truncation and user are accepted and ignored. Caching is automatic.
  • Images go in input_image parts as base64 data URLs. A web link or a file_id is rejected with image_url_not_supported; input_file and input_audio parts are rejected with unsupported_content.

Errors use the OpenAI shape, {"error": {"message", "type", "code", "param"}}, with the same codes as Chat Completions; param names the Responses field. See Error codes.