Rate limits
How Modellane limits request rates per account, what a 429 response means, and how to retry with backoff without losing requests.
Rate limits protect every account from noisy neighbours. When you go over a limit the API answers with HTTP 429 and you can retry after a short wait.
Limits apply per project: every API key of a project shares them. There are two, and a request must pass both.
| Limit | Default | What counts |
|---|---|---|
| Concurrent requests | 20 | Requests in flight at the same moment. A request counts from the time it is accepted until the response is complete or the connection closes. |
| Requests per minute | 300 | Requests accepted in the last 60 seconds, as a sliding window. |
There is no separate token-per-minute limit. If your workload needs more, for example a coding agent that runs several sub-tasks at once, contact us to raise the limits of your project.
When you hit a limit#
The API answers with HTTP 429 and a Retry-After header that gives the wait in seconds. The error code tells you which limit was reached:
| Param | Value |
|---|---|
concurrency_limit | Too many requests in flight. A slot frees up as soon as one of your running requests finishes; Retry-After is 1 second. |
rate_limit_exceeded | Too many requests in the last minute. Retry-After is the time until the oldest request leaves the window. |
model_capacity | The model itself is busy, independent of your own limits. Retry-After is 5 seconds. |
A rejected request is not billed and does not count toward the per-minute window.
Retry with backoff#
Wait at least the Retry-After value, then retry. If the retry is also rejected, double the wait each time and add a little random jitter so parallel workers do not retry in lockstep. The official OpenAI SDKs already retry 429 responses and read Retry-After; you can tune them with max_retries (Python) or maxRetries (Node.js).
Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MODELLANE_API_KEY"],
base_url="https://usemodellane.com/v1",
max_retries=5,
timeout=300,
)Node.js
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.MODELLANE_API_KEY,
baseURL: "https://usemodellane.com/v1",
maxRetries: 5,
timeout: 300_000,
})Long requests and timeouts#
- Without streaming the whole answer must arrive within 120 seconds. If it does not, the request ends with HTTP 504
upstream_timeoutand is not billed. Use streaming for long outputs and reasoning models. - With streaming there is no limit on the total length of a response. The stream must start within 120 seconds, and it ends if the model sends nothing for 120 seconds. While a model is thinking and no text is ready yet, the API may send SSE comment lines (
: ping) to keep the connection open. Standard SSE clients and the OpenAI SDKs ignore them; if you parse the stream yourself, skip lines that start with:. - Client timeouts. Set your HTTP client timeout above these values, especially for reasoning models and long outputs. With
stream: trueyou see progress right away and a slow answer is easier to tell apart from a stuck one. - No automatic retries on our side. If a request fails, we do not resend it to the model, so a retry from your side never produces a hidden duplicate charge.