Logo

Wan 3.0 Text to Video Generator

Overview

What Is Wan 3.0 Text to Video?

Text to video is the simplest way to use Wan 3.0, the video model from Alibaba's Tongyi Lab: you write a prompt and send no image or other media. The model plans the picture, the motion and the sound from your words alone.

One prompt can describe a single shot or a short storyboard. For multi-shot clips Alibaba's guide suggests shots of about 4 to 6 seconds, each with a timestamp, up to 30 seconds in total. Dialogue, ambience and music are generated together with the picture.

Wan3 AI is an independent site, not an Alibaba product. We call Wan 3.0 through a third-party API and show the credit cost of every clip before you generate it.

Text to video on Wan3 AI

Input
A text prompt, up to 5,000 characters
Clip length
2–30 seconds
Resolution
480p, 720p or 1080p
Aspect ratio
16:9, 9:16, 1:1, 4:3, 3:4 or adaptive
Audio
Generated with the video; can be switched off at the same price
Lowest cost
14 credits (5 s at 480p)

Facts checked on 11 October 2026 against Alibaba's documentation.

How it works

Text to Video in Three Steps

  1. 01

    Write the shot

    Say who or what is in frame, where it is, what moves and how the camera follows. Quote any dialogue word for word.

  2. 02

    Set length and format

    Pick 2–30 seconds, 480p to 1080p and an aspect ratio. The credit cost updates before you submit.

  3. 03

    Generate and refine

    Review the clip, change one thing in the prompt at a time, then render the final version at a higher resolution. Failed generations are refunded automatically.

Use cases

What to Make With Text to Video

Story scenes and short drama

Write a scene with characters, actions and spoken lines and get it back as one take of up to 30 seconds.

Concepts and pitches

Try out an ad idea, a storyboard or a mood before you shoot anything or design a single frame.

Social clips

Render vertical 9:16 clips for TikTok, Reels and Shorts straight from a script.

Narrated explainers

Describe a process or a place and the narration you want; voice and picture come out together.

Prompt tips

Write Better Text-to-Video Prompts

01

Subject, scene, motion

Start with who or what is in frame, then where it is, then what moves. Add camera, style and sound once the basic shot works.

02

Time long takes

For longer clips, split the prompt into timed shots of a few seconds each (0–5 s, 5–10 s …) and say how the clip ends.

03

Script the sound

Quote dialogue word for word, name the ambience, and write "no background music" if you want none.

FAQ

Wan 3.0 Text to Video FAQ

Do I need an image for Wan 3.0 text to video?

No. Text to video uses only your prompt. If the clip has to start or end on a specific picture, use Image to Video; if a face, product or outfit has to match your photos, use Reference to Video.

How long can a text-to-video clip be?

Any whole number of seconds from 2 to 30. Longer clips cost more credits and take longer to render.

Does it generate sound?

Yes. Wan 3.0 generates dialogue, ambience and music with the picture. You can switch audio off; the price stays the same.

Which languages can I write prompts in?

Alibaba documents prompts in English and Chinese. Write in one of them for the most predictable results.

How much does it cost?

A 5-second clip costs 14 credits at 480p, 28 at 720p and 55 at 1080p; 30 seconds at 1080p costs 329. The form shows the exact cost for your settings before you generate, and failed generations are refunded automatically.

Write Your First Wan 3.0 Shot

Describe one clear shot, pick a length and check the credit cost before you generate.