Wan 3
Single-shot video from text: duration, resolution, aspect ratio, and audio are each chosen separately.
modellane/wan-3 · Text to video · Alibaba
A video model that turns text into moving scenes. It reads the prompt as camera movement, light, and scene flow, and is strong at keeping objects and characters consistent across a single-shot take. Duration and resolution are chosen from the parameter list in the run form; a soundtrack that fits the scene can be added optionally. It works in horizontal, vertical, and square ratios, so one prompt can yield both an ad and a vertical social cut.
- Vendor
- Alibaba
- Task
- Text to video
- Output
- Video
- Price
- $0.06 – $0.12 · per second
- Longest run
- 10 seconds
Input parameters
- Prompt
prompt— required, at most 20000 characters Describe what you want produced as concretely as you can. - Duration
duration— optional, default: 5, range: 2–10 sn, affects price A longer duration raises the price directly. - Resolution
resolution— optional, default: 480p, options: 480p, 720p, affects price Higher resolution raises the per-second price. - Aspect ratio
ratio— optional, default: 16:9, options: adaptive, 16:9, 9:16, 4:3, 3:4, 1:1 "adaptive" picks the ratio from the source. - Generate audio
audio— optional, advanced, default: true When off, the video is produced silent.
Recommended use cases
- Ad and promo clips — Short single-shot scenes for a product or a place, long enough to use without a cut on the timeline.
- Vertical social media video — Short vertical clips watched in a feed; the chosen ratio is applied without breaking the composition.
- Scene and atmosphere tests — See a shot's light, camera movement, and pace in motion before committing to production.
- Short scenes with audio — Ambient sound is generated with the scene, so a watchable clip comes out without a separate sound design step.
Where it is strong
- Reads camera movement from the prompt and holds it in one take
- Maintains object and character consistency across the scene
- Can generate a soundtrack that matches the picture
- Holds the framing in horizontal, vertical, and square ratios
Limits
- Takes no reference image; use Wan 3 Image to Video to supply a starting frame
- No cuts, transitions, or multi-shot editing; the output is a single take
- Lip sync on talking faces is not reliable
- Not suitable for placing legible text in the picture
Example prompt
Slow dolly shot through a rain-soaked Istanbul side street at night, neon shop signs reflecting in puddles, steam rising from a manhole, cinematic anamorphic lookExample request
POST /v1/queue/modellane/wan-3Playground
The identifier, price, parameters, and example requests on this page can be read without signing in. Running the model from the panel sends a real generation request and charges the project balance; the playground therefore requires an account.