OpenAI Responses API
Use the OpenAI Responses format with Modellane: input items, function tools, streaming events, and what changes for Codex users.
The Responses API is the newer OpenAI request format, used by Codex and recent SDK features. Modellane serves it statelessly on the same models and prices.
Modellane serves the OpenAI Responses API at POST /v1/responses, on the same base URL, keys, models and prices as Chat Completions. It is the format Codex uses, and the one behind client.responses in the OpenAI SDKs. Responses are stateless: nothing is stored, so every request carries the whole conversation.
Make a request#
curl
curl https://usemodellane.com/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MODELLANE_API_KEY" \
-d '{
"model": "lane-1",
"instructions": "You are a helpful assistant.",
"input": "Hello!"
}'Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MODELLANE_API_KEY"],
base_url="https://usemodellane.com/v1",
)
response = client.responses.create(
model="lane-1",
instructions="You are a helpful assistant.",
input="Hello!",
)
print(response.output_text)Node.js
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.MODELLANE_API_KEY,
baseURL: "https://usemodellane.com/v1",
})
const response = await client.responses.create({
model: "lane-1",
instructions: "You are a helpful assistant.",
input: "Hello!",
})
console.log(response.output_text)instructions becomes the system prompt. input is a string for a single prompt, or an array of items for a conversation.
Conversations without stored state#
Because nothing is stored, previous_response_id, conversation, prompt and background: true are rejected with unsupported_parameter, and there is no endpoint to fetch or delete an earlier response. Keep the history yourself and send it in input: user and assistant message items, plus the function_call and function_call_output items of earlier tool calls.
Python
history = [{"role": "user", "content": "Name a prime number."}]
first = client.responses.create(model="lane-1", input=history)
history += [
{"role": "assistant", "content": first.output_text},
{"role": "user", "content": "And the next one?"},
]
second = client.responses.create(model="lane-1", input=history)
print(second.output_text)Function tools#
Declare tools with type: "function", a name and a JSON Schema in parameters. The model replies with function_call items; run the function and send the result back as a function_call_output item with the same call_id, together with the call itself.
Python
tools = [{
"type": "function",
"name": "get_weather",
"description": "Return the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
}]
input_items = [{"role": "user", "content": "What is the weather in Paris?"}]
response = client.responses.create(model="lane-1", input=input_items, tools=tools)
for item in response.output:
if item.type == "function_call":
input_items.append(item)
input_items.append({
"type": "function_call_output",
"call_id": item.call_id,
"output": "Sunny, 22 C",
})
final = client.responses.create(model="lane-1", input=input_items, tools=tools)
print(final.output_text)Custom tools (type: "custom") receive one free-text input. Their grammar is passed to the model as a hint, not enforced. Tools grouped in a namespace are sent to the model under flat names and come back with their namespace. Hosted tools, such as web search, file search, code interpreter and image generation, are removed from the request without an error, because they run on OpenAI's servers.
Streaming#
With stream: true the response arrives as named server-sent events. Text comes in response.output_text.delta, tool arguments in response.function_call_arguments.delta, and the last event is response.completed with the full response and its usage. If the output stops early, response.incomplete replaces it; an error after the stream started arrives as response.failed. The SDKs handle the events for you:
Python
stream = client.responses.create(model="lane-1", input="Count to five.", stream=True)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)The full event list is on Responses.
Reasoning#
Models that reason do so on their own; reasoning.effort and reasoning.summary have no effect. When the model produces reasoning, it comes back as a reasoning item whose summary holds the text, and the tokens count as output tokens. Encrypted reasoning is never returned, and reasoning items sent back in input are dropped.
Notes for Codex#
- Set
model_context_windowin the Codex config. Codex does not read context lengths from a custom provider and otherwise assumes a fixed size; see Codex. - Codex sends its web search tool by default. It is skipped, so the model answers without web access.
- Codex already sends the full history with
store: false, which is what this endpoint needs. - Images from Codex's image viewing tool need a model with Vision; on a text-only model the request fails with
unsupported_content. - Up to 128 tools fit in one request. A setup with many MCP servers can go over that and get a 400 error.
Other fields#
max_output_tokensworks likemax_tokensin Chat Completions; when it is reached, the status isincompletewith reasonmax_output_tokens.text.formattakestext,json_objectorjson_schema, likeresponse_format. See JSON output.metadatais returned unchanged;storeis always reported asfalse.include,parallel_tool_calls,prompt_cache_key,service_tier,truncationanduserare accepted and ignored. Caching is automatic.- Images go in
input_imageparts as base64 data URLs. A web link or afile_idis rejected withimage_url_not_supported;input_fileandinput_audioparts are rejected withunsupported_content.
Errors use the OpenAI shape, {"error": {"message", "type", "code", "param"}}, with the same codes as Chat Completions; param names the Responses field. See Error codes.