Alibaba's newest Wan model generates picture and sound in one pass. 5 second clips at 480p, 720p or 1080p, five aspect ratios, no watermark. Wan 3.0 and Wan 3.0 Prime are both available from the Starter plan on Aura AI.
Wan 3.0 is the third generation of Alibaba's Wan video family, developed by Tongyi Lab and opened in public beta on August 6, 2026 through Alibaba Cloud Model Studio. It follows Wan 2.7 and Wan 2.6 Multi-Shot, which already generated audio with the picture, and takes the same architecture in two directions at once: much longer single-pass clips and a far wider set of inputs.
Upstream, Wan 3.0 produces native 30 second clips in one pass, double the 15 second ceiling of Wan 2.6 and 2.7, with dialogue, music and sound effects generated together with the footage. Its signature feature is Omni-Reference: up to 20 references in a single job, spanning images, video, audio, documents (PDF, Word, slides, spreadsheets, Markdown) and webpage URLs, so the model can build a video from a deck or a product page. Alibaba also reports more expressive faces and micro-expressions, steadier subject references and cleaner on-screen text than 2.7. Unlike the openly released Wan 2.2, Wan 3.0 is a closed, API-only model with no published weights. Wan 3.0 Prime is the accelerated variant, described by Alibaba as significantly faster while keeping output quality.
On Aura AI, Wan 3.0 and Wan 3.0 Prime run in text-to-video and image-to-video mode with 5 second clips at 480p, 720p or 1080p in five aspect ratios from 16:9 to 9:16. Output is delivered without watermark and ready for commercial use. Both variants sit in the same studio as Seedance 2.5, Kling 3.0 and the rest of the Wan line, so you can switch between models without leaving the tool.
Six capabilities that separate 3.0 from Wan 2.7 and Wan 2.6
Dialogue, music and ambience arrive in the same pass as the footage. A clip comes back scored rather than silent, so a spoken line or a beat cue in the prompt is already aligned with the motion.
Alibaba's model generates native 30 second clips in one pass, twice what Wan 2.7 allowed. On Aura AI, Wan 3.0 runs at 5 seconds per generation, which keeps costs predictable while you use the same model quality.
Upstream, Wan 3.0 accepts up to 20 references per job: images, video, audio, documents and webpages. The model can turn a slide deck, a PDF or a product page into a video sequence, a first for the Wan line.
Alibaba reports better micro-expressions and steadier subject identity than Wan 2.7. Close-ups on people, reaction shots and short dialogue scenes hold up with fewer re-rolls.
Signage, labels, UI mockups and captions render more legibly than in earlier Wan releases. Useful for product shots and explainer clips where a readable word on screen matters.
Prime is the accelerated version of Wan 3.0, built for faster turnaround with the same quality target. On Aura AI it shares every setting with standard Wan 3.0 and costs one credit more per clip.
Four steps from prompt to a finished clip with sound
Cost = 6 credits base (7 for Prime) for a 5 second clip, multiplied by the resolution factor
| Model, duration | 480p (x0.6) | 720p (x0.8) | 1080p (x1) |
|---|---|---|---|
| Wan 3.0, 5 seconds | 4 credits | 5 credits | 6 credits |
| Wan 3.0 Prime, 5 seconds | 5 credits | 6 credits | 7 credits |
How the new Alibaba release compares to its siblings on Aura AI
| Feature | Wan 3.0 | Wan 2.7 | Wan 2.6 Multi-Shot | Seedance 2.5 |
|---|---|---|---|---|
| Max resolution on Aura AI | 1080p | 1080p | 1080p | 1080p |
| Duration on Aura AI | 5 s | 2 to 15 s | 5, 10 or 15 s | 4 to 15 s |
| Max duration upstream | 30 s | 15 s | 15 s | 15 s |
| Native audio | ✓ Single pass | ✓ Yes | ✓ Yes | ✓ Single pass, 10+ languages |
| Reference inputs upstream | ✓ Up to 20, incl. documents and webpages | Image + text | Image + text | ✓ Up to 50 |
| Aspect ratios | 5 | 5 | 5 | 6 |
| Credits, 5 s at 1080p | 6 (Prime 7) | 5 | 4 | 13 |
| Plan | Starter+ | Starter+ | All plans | Starter+ |
Where single-pass audio and better faces pay off in a 5 second clip
A 9:16 clip with the product in motion and a sound bed already mixed. Use image-to-video with a packshot to lock the exact design, then let Wan 3.0 add the camera move and the audio.
Wan 3.0's improved micro-expressions make a single spoken line believable. Ideal for reaction shots, testimonial-style snippets and short character moments in ads.
Cleaner on-screen text means a label, a price tag or a UI element stays readable. Pair a still of your interface with image-to-video for a quick animated walkthrough.
Atmospheric establishing shots with matching ambience: rain, traffic, wind, a distant crowd. Generate several 5 second clips and cut them together in the video editor.
Common questions about Wan 3.0 on Aura AI
Video and audio in one pass, up to 1080p, no watermark. Wan 3.0 and Wan 3.0 Prime, available from the Starter plan, worldwide.