Production-ready models

Image, video, and audio generation models in the Modellane catalog, each listed with its vendor, task, unit price, and input parameters.

Every identifier here is used directly as the model name in the queue API; the provider version behind it may change, the identifier does not.

Text to image

Seedream 4.5

modellane/seedream-4.5 · ByteDance

An image model that turns text into high-resolution, photorealistic frames. It is strong at texture, skin, and material detail; at building natural light and depth; and at rendering legible lettering and typography inside the image. It holds a composition together across wide aspect ratios, which is why it is the pick for horizontal work such as cover images and banners.

Vendor
ByteDance
Task
Text to image
Output
Image
Price
$0.06 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Size size — optional, default: 2048*2048, options: 2048*2048, 2560*1440, 1440*2560, 2304*1728, 1728*2304

Recommended use cases

  • Product and editorial photography — Studio-lit product frames for catalogs, landing pages, and decks; material and surface detail are preserved.
  • Images with text and typography — Short headlines set legibly inside posters, covers, and social media images.
  • Wide compositions — Full-frame output for wide and tall ratios: web hero images, banners, and video covers.
  • Concept and pre-production — Compare location, set, and atmosphere options quickly at high resolution.

Qwen Image 3

modellane/qwen-image-3 · Alibaba

A multilingual text-to-image model. It understands Turkish prompts directly and can place short Turkish, English, or Chinese text inside the image. It returns up to four variations per request and supports a negative prompt for keeping unwanted elements out, which suits fast iteration rounds and bulk variation work.

Vendor
Alibaba
Task
Text to image
Output
Image
Price
$0.045 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Negative prompt negative_prompt — optional, advanced, at most 2000 characters Things you don't want to see in the output.
  • Size size — optional, default: 1024*1024, options: 1024*1024, 2048*2048, 1792*1024, 1024*1792, 1536*1152, 1152*1536
  • Number of images n — optional, default: 1, range: 1–4, affects price

Recommended use cases

  • Generating from Turkish prompts — Generate an image by describing it in Turkish, without translating the prompt into English first.
  • Variation sweep — Pull several options from one prompt in a single pass and pick a direction; it shortens iteration rounds.
  • Social media images — Share images in square, vertical, and horizontal ratios, plus card designs with short headlines.
  • Blog and content images — Prepare explanatory scenes to sit alongside an article, simplified with a negative prompt.

Nano Banana 2

modellane/nano-banana-2 · Google

A fast image model built on Google Gemini 3.1 Flash. Strong in figurine, character, and stylised art styles; it produces fluid, striking images within seconds. With ten aspect ratios and resolution up to 4K, it fits everything from posters to social posts.

Vendor
Google
Task
Text to image
Output
Image
Price
$0.1 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Aspect ratio aspect_ratio — optional, default: 1:1, options: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
  • Output format output_format — optional, advanced, default: png, options: png, jpeg
  • Seed seed — optional, advanced, default: -1 The same seed and the same settings produce a similar result; leave it empty and one is picked at random.

Recommended use cases

  • Figurine and character design — Collectible figurine, game character, or mascot concepts; clean and shareable in a 3D render aesthetic.
  • Stylised social media images — Eye-catching share images in square, vertical, and cinematic ratios.
  • Fast concept tests — Turn an idea into a concrete image within seconds and pick a direction.
  • Posters and cover images — One model for the full ratio range, from a 21:9 cinematic crop to a 9:16 vertical poster.

Seedream 5.0 Lite

modellane/seedream-5.0-lite · ByteDance

ByteDance's reasoning-first image model. It produces images grounded in current information through real-time web search and chain-of-thought, and keeps place, physics, and light consistent. Built-in editing support lets you do style transfer and multi-source composition on the same model.

Vendor
ByteDance
Task
Text to image
Output
Image
Price
$0.04 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Size size — optional, default: 1024*1024, options: 1024*1024, 2048*2048, 1792*1024, 1024*1792

Recommended use cases

  • Images grounded in current information — Visualise current places, products, and events through web search.
  • Location and atmosphere design — Produce architectural, interior, and set concepts with physical consistency.
  • Style transfer and editing — Carry the style of an existing image onto another scene; combine several images in one frame.
  • Typography and posters — Strong text handling for in-image copy, headlines, and poster design.

Grok Imagine 2

modellane/grok-imagine-2 · xAI

xAI's fast and inexpensive image model. It returns up to four variations at a time; with 14 aspect ratios and 1K-2K resolution support it fits everything from posters to profile pictures. Its low unit cost makes it ideal for high-volume production and bulk variation sweeps.

Vendor
xAI
Task
Text to image
Output
Image
Price
$0.025 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Number of images n — optional, default: 1, range: 1–4, affects price

Recommended use cases

  • Bulk variation sweep — Get four options from the same prompt in one pass and pick a direction; it shortens iteration rounds.
  • High-volume production — Produce large numbers of images at a low unit cost; catalogs and content calendars.
  • Social media formats — One model for every platform with 14 ratios: square, vertical, horizontal, profile.
  • Fast prototypes and drafts — Turn an idea into a concrete image within seconds and show it to stakeholders.

Flux 2 Pro

modellane/flux-2-pro · Black Forest Labs

The flagship of the Black Forest Labs Flux 2 family. It produces studio-quality images and is strong at sharp detail, natural light, and legible text handling. There is no parameter clutter — one good prompt is enough. Ideal for posters, product shots, and corporate identity work.

Vendor
Black Forest Labs
Task
Text to image
Output
Image
Price
$0.04 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Number of images n — optional, default: 1, range: 1–4, affects price

Recommended use cases

  • Corporate and editorial images — Studio-quality frames for brand identity, catalogs, and print work.
  • Poster design — Posters, covers, and announcement images with legible text and typography inside the frame.
  • Product photography — Product frames with texture, material, and light detail; no studio setup needed.
  • Fast variation production — Pick the best frame from up to four variations in one pass.

Flux 2 Turbo

modellane/flux-2-turbo · Black Forest Labs

The speed-optimised variant of the Flux 2 family. Designed for low latency and high throughput, it is ideal for real-time production pipelines and batch work. It produces results close to Pro quality much faster and more cheaply.

Vendor
Black Forest Labs
Task
Text to image
Output
Image
Price
$0.015 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Number of images n — optional, default: 1, range: 1–4, affects price

Recommended use cases

  • Bulk content production — High-volume image production for blogs, e-commerce, and content calendars.
  • Real-time applications — Low-latency generation inside interactive applications where the user is waiting.
  • Rapid prototyping — Turn an idea into an image within seconds and pick a direction; it shortens the trial-and-error loop.
  • Social media series — Produce many images on one theme quickly and publish them on a schedule.

Flux 2 Flash

modellane/flux-2-flash · Black Forest Labs

The fastest and least expensive variant of the Flux 2 family. Optimised for lightning-fast production of posters, logos, product shots, and social media content, it is the first choice for volume production and budget-conscious work.

Vendor
Black Forest Labs
Task
Text to image
Output
Image
Price
$0.01 · per image

Input parameters

  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.
  • Number of images n — optional, default: 1, range: 1–4, affects price

Recommended use cases

  • High-volume production — Thousands of images at the lowest unit cost; catalogs, listings, and feeds.
  • Social media images — Produce daily posts, stories, and cover images within seconds.
  • Logo and icon tests — Fast drafts for brand identity, logo variations, and icon sets.
  • Drafts and first pass — Draft with Flash and finish with Pro; it lowers the cost.

Text to video

Wan 3

modellane/wan-3 · Alibaba

A video model that turns text into moving scenes. It reads the prompt as camera movement, light, and scene flow, and is strong at keeping objects and characters consistent across a single-shot take. Duration and resolution are chosen from the parameter list in the run form; a soundtrack that fits the scene can be added optionally. It works in horizontal, vertical, and square ratios, so one prompt can yield both an ad and a vertical social cut.

Vendor
Alibaba
Task
Text to video
Output
Video
Price
$0.06 – $0.12 · per second
Longest run
10 seconds

Input parameters

  • Prompt prompt — required, at most 20000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 2–10 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 480p, options: 480p, 720p, affects price Higher resolution raises the per-second price.
  • Aspect ratio ratio — optional, default: 16:9, options: adaptive, 16:9, 9:16, 4:3, 3:4, 1:1 "adaptive" picks the ratio from the source.
  • Generate audio audio — optional, advanced, default: true When off, the video is produced silent.

Recommended use cases

  • Ad and promo clips — Short single-shot scenes for a product or a place, long enough to use without a cut on the timeline.
  • Vertical social media video — Short vertical clips watched in a feed; the chosen ratio is applied without breaking the composition.
  • Scene and atmosphere tests — See a shot's light, camera movement, and pace in motion before committing to production.
  • Short scenes with audio — Ambient sound is generated with the scene, so a watchable clip comes out without a separate sound design step.

Seedance 2.5

modellane/seedance-2.5 · ByteDance

ByteDance's newest video model. It produces up to 30 seconds of 4K video in one pass and keeps characters, products, and places consistent across scenes with up to 50 multimodal references. Audio is generated along with the picture; with multilingual subtitles and in-frame text support it is a complete production pipeline.

Vendor
ByteDance
Task
Text to video
Output
Video
Price
$0.168 · per second
Longest run
30 seconds

Input parameters

  • Prompt prompt — required, at most 20000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 4–30 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 720p, options: 480p, 720p, 1080p Higher resolution raises the per-second price.
  • Aspect ratio ratio — optional, default: adaptive, options: adaptive, 16:9, 9:16, 1:1, 4:3, 3:4 "adaptive" picks the ratio from the source.
  • Generate audio generate_audio — optional, advanced, default: true When off, the video is produced silent.
  • Output format output_format — optional, advanced, default: mp4, options: mp4, mov

Recommended use cases

  • Long single-shot scenes — Up to 30 seconds of uninterrupted footage; ideal for short films, ads, and promos.
  • Character and brand consistency — Keep the same character, product, and place across scenes with 50 references.
  • Multilingual content production — Video in 10+ languages with in-frame subtitles and text; no localisation step needed.
  • Cinematic production — Broadcast-ready output with 4K resolution, 10-bit colour, and synchronised audio.

H3 Max

modellane/h3-max · MiniMax

MiniMax's cinematic video model. It produces 5-15 second, 24fps videos with audio from text. With 480p and 768p resolution options, six aspect ratios, and optional prompt expansion it fits everything from social media to ad clips. It can produce a 5-second clip in roughly 3 seconds — which speeds the jump from idea to result.

Vendor
MiniMax
Task
Text to video
Output
Video
Price
$0.06 · per second
Longest run
15 seconds

Input parameters

  • Prompt prompt — required, at most 7000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 5–15 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 768p, options: 480p, 768p Higher resolution raises the per-second price.
  • Aspect ratio ratio — optional, default: 16:9, options: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 "adaptive" picks the ratio from the source.
  • Prompt expansion prompt_expansion — optional, advanced, default: false When on, the model enriches the prompt with its own rules.

Recommended use cases

  • Social media ad clips — Product or place promos in short, vertical, and square formats.
  • Cinematic scene tests — See a scene's light, camera, and pace before committing to production.
  • Fast draft production — A 5-second clip in 3 seconds; test an idea in motion right away.
  • Multi-format delivery — Produce 21:9 cinematic, 16:9 horizontal, and 9:16 vertical from the same prompt.

H3 Fast

modellane/h3-fast · MiniMax

The fastest and least expensive variant in the H3 family. It produces 5-15 second, 24fps videos with audio at 480p. It is tuned for high-volume draft production, bulk variation sweeps, and quick tests at the idea stage. An ideal first step: turn an idea into motion within seconds, then finish with H3 Max or Seedance 2.5.

Vendor
MiniMax
Task
Text to video
Output
Video
Price
$0.055 · per second
Longest run
15 seconds

Input parameters

  • Prompt prompt — required, at most 7000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 5–15 sn, affects price A longer duration raises the price directly.
  • Aspect ratio ratio — optional, default: 16:9, options: 16:9, 9:16, 1:1, adaptive "adaptive" picks the ratio from the source.
  • Prompt expansion prompt_expansion — optional, advanced, default: false When on, the model enriches the prompt with its own rules.

Recommended use cases

  • Bulk draft production — Produce dozens of variations on an idea within seconds and cut them down.
  • Idea-stage tests — Test the scene in motion before committing to production.
  • Budget-friendly content — Produce high-volume social media content at a low unit cost.
  • First pass, then upgrade — Draft with H3 Fast and finish with H3 Max or Seedance 2.5.

Image to video

Wan 3 Görselden Video

modellane/wan-3-i2v · Alibaba

A video model that sets an uploaded frame in motion. The image becomes the first frame and the prompt describes what happens in it: how the camera moves, where the elements go, and how the scene develops. The source image's composition, colours, and characters are preserved, so with a ready image in hand the result is far more predictable than generating from scratch. Duration and resolution are chosen from the parameter list in the run form; audio is optional.

Vendor
Alibaba
Task
Image to video
Output
Video
Price
$0.06 – $0.12 · per second
Longest run
10 seconds

Input parameters

  • Source image image — required The frame the generation starts from.
  • Prompt prompt — required, at most 20000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 2–10 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 480p, options: 480p, 720p, affects price Higher resolution raises the per-second price.
  • Generate audio audio — optional, advanced, default: true When off, the video is produced silent.

Recommended use cases

  • Animating a still image — Turn a product photo, illustration, or poster you already have into short motion; the composition stays exactly as it is.
  • Video version of a campaign image — Derive a short clip for the feed from an approved campaign frame without disturbing its framing.
  • Character and brand consistency — Carry over the elements that must stay recognisable, such as a face, a garment, or packaging; text-to-image changes them every time.
  • Camera movement tests — Try different camera moves on the same frame and see which one suits the scene.

Seedance 2.5 Görselden Video

modellane/seedance-2.5-i2v · ByteDance

Sets an uploaded frame in motion with the Seedance 2.5 engine for up to 30 seconds. The source image's composition, colours, and characters are preserved; the prompt only describes the motion. Unlike the full 50-reference version it works from a single frame, so with a ready image in hand the result is far more predictable than generating from scratch.

Vendor
ByteDance
Task
Image to video
Output
Video
Price
$0.168 · per second
Longest run
30 seconds

Input parameters

  • Source image image — required The frame the generation starts from.
  • Prompt prompt — required, at most 20000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 4–30 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 720p, options: 480p, 720p, 1080p Higher resolution raises the per-second price.
  • Generate audio generate_audio — optional, advanced, default: true When off, the video is produced silent.
  • Output format output_format — optional, advanced, default: mp4, options: mp4, mov

Recommended use cases

  • Animating a still image — Turn a product photo, illustration, or poster into short motion.
  • Video version of a campaign image — Derive video from an approved campaign frame without disturbing its framing.
  • Character and brand consistency — Carry elements such as a face, a garment, or packaging over from the source frame.
  • Camera movement tests — Try different camera moves on the same frame and choose the one you want.

H3 Max Görselden Video

modellane/h3-max-i2v · MiniMax

Sets an uploaded frame in motion with the H3 Max engine for up to 15 seconds. The first frame is preserved, and an optional end frame decides where the motion stops. The prompt describes the movement; the model carries the composition, the colours, and the characters over from the source frame.

Vendor
MiniMax
Task
Image to video
Output
Video
Price
$0.06 · per second
Longest run
15 seconds

Input parameters

  • Source image image — required The frame the generation starts from.
  • Prompt prompt — required, at most 7000 characters Describe what you want produced as concretely as you can.
  • Duration duration — optional, default: 5, range: 5–15 sn, affects price A longer duration raises the price directly.
  • Resolution resolution — optional, default: 768p, options: 480p, 768p Higher resolution raises the per-second price.
  • Prompt expansion prompt_expansion — optional, advanced, default: false When on, the model enriches the prompt with its own rules.

Recommended use cases

  • Animating a product image — Turn a static product photo into a clip that rotates, pushes in, or catches the light.
  • First and last frame control — Supply the opening and closing frames and leave the motion in between to the model.
  • Brand consistency — Carry over the elements that must stay recognisable, such as packaging, a logo, or a product.
  • Fast animation drafts — Set an illustration or concept drawing in motion within seconds.

Text to audio

MiniMax Music 3

modellane/minimax-music-3 · MiniMax

A model that turns text into a finished music track. The prompt describes genre, mood, tempo, instruments, and vocal colour; the model produces the arrangement, the instrumentation, and the mix together. Lyrics that are supplied get sung; if none are supplied, the model writes its own. With the instrumental option on, the track comes out without vocals. The output is a finished audio file whose format is chosen from the parameter list in the run form — so the track can go straight into an edit.

Vendor
MiniMax
Task
Text to audio
Output
Audio
Price
$0.22 · per run

Input parameters

  • Prompt prompt — required, at most 2000 characters Describe what you want produced as concretely as you can.
  • Instrumental is_instrumental — optional, default: false When on, the track is produced without vocals and any lyrics are ignored.
  • Let the model write the lyrics lyrics_optimizer — optional, advanced, default: true When on, the lyrics are derived from the prompt; the lyrics field must stay empty.
  • Lyrics lyrics — optional, advanced, at most 3500 characters Line breaks separate the lines. If you are writing lyrics, turn "Let the model write the lyrics" off.
  • File format format — optional, advanced, default: mp3, options: mp3, wav

Recommended use cases

  • Music for video and ads — An original bed described to match the clip's length and tempo; no browsing stock libraries.
  • Game and app audio — Produce loopable tracks with their own mood for each menu, level, or screen.
  • Lyric demo — Hear written lyrics in a vocal arrangement; they are typed into the lyrics field line by line.
  • Podcast theme — A short, recognisable theme for the intro and outro; with the instrumental option it can also run under speech.

Image to image

Flux 2 Pro Edit

modellane/flux-2-pro-edit · Black Forest Labs

A Flux 2 Pro variant that edits an existing image from a natural-language instruction. It offers colour changes, style transfer, background replacement, adding and removing objects, and precise control with hex colour codes. The source image's composition is preserved and only the described change is applied.

Vendor
Black Forest Labs
Task
Image to image
Output
Image
Price
$0.08 · per image

Input parameters

  • Source image image — required The frame the generation starts from.
  • Prompt prompt — required, at most 5000 characters Describe what you want produced as concretely as you can.

Recommended use cases

  • Background replacement — Replace a product photo's background with a studio, nature, or a location.
  • Style transfer — Turn a photograph into oil paint, comic, watercolour, or pixel art.
  • Product colour variations — Derive different colour options of the same product from a single frame.
  • Adding and removing objects — Add new elements to an image or clean up unwanted objects.