Story scenes and short drama
Write a scene with characters, actions and spoken lines and get it back as one take of up to 30 seconds.
Wan3 AIDescribe a shot and Wan 3.0 renders 2–30 seconds at up to 1080p, with sound. No image needed.
Text to video is the simplest way to use Wan 3.0, the video model from Alibaba's Tongyi Lab: you write a prompt and send no image or other media. The model plans the picture, the motion and the sound from your words alone.
One prompt can describe a single shot or a short storyboard. For multi-shot clips Alibaba's guide suggests shots of about 4 to 6 seconds, each with a timestamp, up to 30 seconds in total. Dialogue, ambience and music are generated together with the picture.
Wan3 AI is an independent site, not an Alibaba product. We call Wan 3.0 through a third-party API and show the credit cost of every clip before you generate it.
Facts checked on 11 October 2026 against Alibaba's documentation.
The generator is at the top of this page, with Text to Video already selected.
Say who or what is in frame, where it is, what moves and how the camera follows. Quote any dialogue word for word.
Pick 2–30 seconds, 480p to 1080p and an aspect ratio. The credit cost updates before you submit.
Review the clip, change one thing in the prompt at a time, then render the final version at a higher resolution. Failed generations are refunded automatically.
Write a scene with characters, actions and spoken lines and get it back as one take of up to 30 seconds.
Try out an ad idea, a storyboard or a mood before you shoot anything or design a single frame.
Render vertical 9:16 clips for TikTok, Reels and Shorts straight from a script.
Describe a process or a place and the narration you want; voice and picture come out together.
Habits based on Alibaba's prompt guidance for Wan 3.0.
Start with who or what is in frame, then where it is, then what moves. Add camera, style and sound once the basic shot works.
For longer clips, split the prompt into timed shots of a few seconds each (0–5 s, 5–10 s …) and say how the clip ends.
Quote dialogue word for word, name the ambience, and write "no background music" if you want none.
No. Text to video uses only your prompt. If the clip has to start or end on a specific picture, use Image to Video; if a face, product or outfit has to match your photos, use Reference to Video.
Any whole number of seconds from 2 to 30. Longer clips cost more credits and take longer to render.
Yes. Wan 3.0 generates dialogue, ambience and music with the picture. You can switch audio off; the price stays the same.
Alibaba documents prompts in English and Chinese. Write in one of them for the most predictable results.
A 5-second clip costs 14 credits at 480p, 28 at 720p and 55 at 1080p; 30 seconds at 1080p costs 329. The form shows the exact cost for your settings before you generate, and failed generations are refunded automatically.
Animate a photo as the first frame, with an optional last frame.
Open Image to VideoKeep faces, products and outfits from up to 9 images consistent in a new scene.
Open Reference to VideoEvery Wan model in one place, including Wan 2.6.
Open the generatorBrowse prompts with preview clips and open any of them in the generator.
Browse promptsDescribe one clear shot, pick a length and check the credit cost before you generate.