Vision input

Send images to models that accept them as base64 data URLs inside chat messages, with the supported formats and limits.

Models with image input read pictures sent inside a user message as base64 data URLs. This guide shows the message shape, formats and limits.

Models with image input can describe photos, read screenshots and answer questions about charts. Only some models accept images: they are marked with vision on Models & pricing, and the examples below use lane-1. Sending an image to a text-only model returns 400 unsupported_content.

Message format#

Images go inside a user message. Instead of a string, content becomes an array of parts: text parts and image_url parts. The image itself is embedded in the request as a base64 data URL.

JSON

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What is in this picture?"},
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..."}}
  ]
}

Send an image from a file#

curl

IMAGE=$(base64 < photo.jpg | tr -d '\n')

curl https://usemodellane.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MODELLANE_API_KEY" \
  -d @- <<EOF
{
  "model": "lane-1",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this picture?"},
      {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,$IMAGE"}}
    ]
  }]
}
EOF

Python

import base64
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MODELLANE_API_KEY"],
    base_url="https://usemodellane.com/v1",
)

with open("photo.jpg", "rb") as f:
    image = base64.b64encode(f.read()).decode("ascii")

response = client.chat.completions.create(
    model="lane-1",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is in this picture?"},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
            ],
        }
    ],
)

print(response.choices[0].message.content)

Node.js

import { readFileSync } from "node:fs"
import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.MODELLANE_API_KEY,
  baseURL: "https://usemodellane.com/v1",
})

const image = readFileSync("photo.jpg").toString("base64")

const response = await client.chat.completions.create({
  model: "lane-1",
  messages: [
    {
      role: "user",
      content: [
        { type: "text", text: "What is in this picture?" },
        { type: "image_url", image_url: { url: `data:image/jpeg;base64,${image}` } },
      ],
    },
  ],
})

console.log(response.choices[0].message.content)

Match the media type in the data URL to the file: image/png, image/jpeg, image/webp or image/gif.

Limits#

ParamValue
FormatsPNG, JPEG, WebP and GIF, as data:image/<type>;base64,...
Size per image5 MB after base64 decoding
Images per request16
Request body8 MB in total, including the base64 text

Base64 makes a file about a third larger, so a 5 MB image takes close to 7 MB of the request body. Resize large photos before you send them: smaller images upload faster and usually use fewer input tokens.

Tips#

  • Images count as input tokens. Each image adds to prompt_tokens, billed at the model's input price. Several images in one request multiply that.
  • Put the question next to the image. A text part that says what to look for gives better answers than an image alone.
  • Images in follow-up turns are sent again. The API is stateless, so an image stays in the history you resend. Drop it from later requests once the model no longer needs it. See Multi-turn conversations.