Streaming responses
Stream chat completions token by token with server-sent events, read the final usage chunk, and handle client disconnects.
Set stream to true and the API sends the answer as server-sent events while it is generated, ending with a usage chunk and a done marker.
Without streaming, the API waits until the whole reply is ready and returns it in one response. With stream: true you see the first words within moments, which matters for chat interfaces and for models that think for a while before they answer.
Turn on streaming#
Set stream to true. Add stream_options: {"include_usage": true} if you want the token counts of the request at the end of the stream.
curl
curl -N https://usemodellane.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MODELLANE_API_KEY" \
-d '{
"model": "lane-1",
"messages": [{"role": "user", "content": "Write a haiku about the sea."}],
"stream": true,
"stream_options": {"include_usage": true}
}'Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MODELLANE_API_KEY"],
base_url="https://usemodellane.com/v1",
)
stream = client.chat.completions.create(
model="lane-1",
messages=[{"role": "user", "content": "Write a haiku about the sea."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)Node.js
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.MODELLANE_API_KEY,
baseURL: "https://usemodellane.com/v1",
})
const stream = await client.chat.completions.create({
model: "lane-1",
messages: [{ role: "user", content: "Write a haiku about the sea." }],
stream: true,
stream_options: { include_usage: true },
})
for await (const chunk of stream) {
const text = chunk.choices[0]?.delta?.content
if (text) process.stdout.write(text)
if (chunk.usage) console.log("\n", chunk.usage)
}Event format#
The response has the content type text/event-stream. Each event is a line that starts with data: followed by one JSON chunk, and events are separated by a blank line:
SSE
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"role":"assistant","content":"Waves"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{"content":" fold"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1767225600,"model":"lane-1","choices":[],"usage":{"prompt_tokens":15,"completion_tokens":19,"total_tokens":34,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}}}
data: [DONE]- Join the
delta.contentpieces in order to build the reply. - The chunk that ends the reply carries
finish_reason:stop,lengthortool_calls. - On models that reason, chunks may carry
delta.reasoning_contentbefore the answer starts. See Reasoning output. - Tool calls arrive in pieces under
delta.tool_calls; see Tool calls. data: [DONE]is always the last event of a successful stream.
Keep-alive comments#
While a model is thinking, the stream may stay quiet for a while. The API sends comment lines that start with a colon (: ping) to keep the connection open. Server-sent event parsers skip them; if you parse the stream yourself, ignore any line that does not start with data: .
The usage chunk#
With include_usage set, the event just before data: [DONE] holds an empty choices array and a usage object with the prompt, cached, completion and reasoning token counts of the whole request. These are the numbers the request is billed for. Without include_usage the chunk is not sent, but the request is billed the same way and its usage appears in your dashboard.
Errors during a stream#
Errors found before generation starts (an invalid body, a missing key, a low balance) come back as a normal JSON error with an HTTP status. Once the stream has started the status is already 200, so a failure is reported as one final event that holds an error object instead of a chunk:
SSE
data: {"error":{"message":"The model failed to produce a response.","type":"api_error","code":"upstream_error"}}The stream then ends without data: [DONE]. Treat a stream that ends without [DONE] as incomplete.
Stopping a stream early#
Close the connection to stop generation, for example when a user presses a stop button. In the SDKs, break out of the loop and close the stream (stream.close() in Python, stream.controller.abort() in Node.js). The API stops the request to the model as soon as it sees the disconnect.
Tips#
- Use streaming for anything a person waits on. For batch jobs, the non-streaming response is simpler to handle.
- Set a client read timeout that tolerates long pauses on reasoning models; the keep-alive comments arrive between pauses.
- Log the final
usagechunk if you track cost per request in your own system.