Rate limits

How Modellane limits request rates per account, what a 429 response means, and how to retry with backoff without losing requests.

Rate limits protect every account from noisy neighbours. When you go over a limit the API answers with HTTP 429 and you can retry after a short wait.

Limits apply per project: every API key of a project shares them. There are two, and a request must pass both.

LimitDefaultWhat counts
Concurrent requests20Requests in flight at the same moment. A request counts from the time it is accepted until the response is complete or the connection closes.
Requests per minute300Requests accepted in the last 60 seconds, as a sliding window.

There is no separate token-per-minute limit. If your workload needs more, for example a coding agent that runs several sub-tasks at once, contact us to raise the limits of your project.

When you hit a limit#

The API answers with HTTP 429 and a Retry-After header that gives the wait in seconds. The error code tells you which limit was reached:

ParamValue
concurrency_limitToo many requests in flight. A slot frees up as soon as one of your running requests finishes; Retry-After is 1 second.
rate_limit_exceededToo many requests in the last minute. Retry-After is the time until the oldest request leaves the window.
model_capacityThe model itself is busy, independent of your own limits. Retry-After is 5 seconds.

A rejected request is not billed and does not count toward the per-minute window.

Retry with backoff#

Wait at least the Retry-After value, then retry. If the retry is also rejected, double the wait each time and add a little random jitter so parallel workers do not retry in lockstep. The official OpenAI SDKs already retry 429 responses and read Retry-After; you can tune them with max_retries (Python) or maxRetries (Node.js).

Python

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
    max_retries=5,
    timeout=300,
)

Node.js

import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
  maxRetries: 5,
  timeout: 300_000,
})

Long requests and timeouts#

  • Without streaming the whole answer must arrive within 120 seconds. If it does not, the request ends with HTTP 504 upstream_timeout and is not billed. Use streaming for long outputs and reasoning models.
  • With streaming there is no limit on the total length of a response. The stream must start within 120 seconds, and it ends if the model sends nothing for 120 seconds. While a model is thinking and no text is ready yet, the API may send SSE comment lines (: ping) to keep the connection open. Standard SSE clients and the OpenAI SDKs ignore them; if you parse the stream yourself, skip lines that start with :.
  • Client timeouts. Set your HTTP client timeout above these values, especially for reasoning models and long outputs. With stream: true you see progress right away and a slow answer is easier to tell apart from a stuck one.
  • No automatic retries on our side. If a request fails, we do not resend it to the model, so a retry from your side never produces a hidden duplicate charge.