Multi-turn conversations

The chat API is stateless: learn how to keep a conversation going by sending the message history with every request.

The API does not remember earlier requests. To continue a conversation, send the previous messages again together with the new one.

The chat completions endpoint keeps no session on the server. Each request is answered only from the messages you send with it. To hold a conversation, your code stores the history and sends it again with every new user message.

How it works#

  1. Send the first user message.
  2. Append the assistant's reply to your message list.
  3. Append the next user message and send the whole list.

Python

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
)

messages = [
    {"role": "system", "content": "You are a concise travel assistant."},
    {"role": "user", "content": "What is the highest mountain in Japan?"},
]

# Round 1
response = client.chat.completions.create(model="lane-1", messages=messages)
messages.append({"role": "assistant", "content": response.choices[0].message.content})
print(messages[-1]["content"])

# Round 2
messages.append({"role": "user", "content": "When is the climbing season?"})
response = client.chat.completions.create(model="lane-1", messages=messages)
messages.append({"role": "assistant", "content": response.choices[0].message.content})
print(messages[-1]["content"])

Node.js

import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
})

const messages = [
  { role: "system", content: "You are a concise travel assistant." },
  { role: "user", content: "What is the highest mountain in Japan?" },
]

// Round 1
let response = await client.chat.completions.create({ model: "lane-1", messages })
messages.push({ role: "assistant", content: response.choices[0].message.content })

// Round 2
messages.push({ role: "user", content: "When is the climbing season?" })
response = await client.chat.completions.create({ model: "lane-1", messages })
messages.push({ role: "assistant", content: response.choices[0].message.content })
console.log(messages.at(-1).content)

In the second request the model receives all four messages:

JSON

[
  {"role": "system", "content": "You are a concise travel assistant."},
  {"role": "user", "content": "What is the highest mountain in Japan?"},
  {"role": "assistant", "content": "Mount Fuji, at 3,776 meters."},
  {"role": "user", "content": "When is the climbing season?"}
]

Cost grows with the history#

Every request is billed for all the prompt tokens it sends, so the earlier turns are paid for again in each round. A long chat costs more per message than a short one, and it eventually reaches the model's context length, at which point the API answers with context_length_exceeded.

Two things keep this under control:

  • Context caching. When a request starts with the same messages as an earlier one, that shared start is usually read from cache and billed at the lower cached input price. Appending to the end of the list, as above, keeps the start identical. See Context caching.
  • Trimming. For long sessions, drop or summarize the oldest turns and keep the system message. Trim in larger steps rather than one message per round: every change to the start of the list makes the next request miss the cache.

Reasoning output#

Models marked with reasoning on Models & pricing think before they answer. The thinking comes back in a separate field, reasoning_content, next to content in the message (and in stream deltas):

JSON

{
  "role": "assistant",
  "reasoning_content": "The user asks for the season. The official season is early July to early September...",
  "content": "The official climbing season runs from early July to early September."
}
  • Reasoning tokens are part of completion_tokens and are billed as output. usage.completion_tokens_details.reasoning_tokens shows how many there were.
  • When you continue the conversation, append only content as the assistant message. You do not need to send reasoning_content back; the API ignores it in request messages.
  • Reasoning counts against max_tokens. With a small limit, a model may spend the whole budget thinking and return an empty content with finish_reason: "length". Leave a generous limit, or omit max_tokens to use the model's default.