Shengshu's Vidu Q3 generates picture and sound in one pass, with multilingual lip-sync and camera control. 5 second clips from 360p to 1080p, five aspect ratios, no watermark. Available from the Starter plan on Aura AI, with Vidu Q2 Reference for reference-to-video on every plan.
Vidu Q3 is the flagship video model of Shengshu Technology, the Beijing company that grew out of Tsinghua University's TSAIL lab and runs the Vidu platform. Shengshu unveiled Q3 on January 30, 2026 during its Global Creativity Week in Singapore, presenting it as the first long-form video model to deliver native audio and video in a single output. Upstream it generates up to 16 seconds of synchronized sound and picture at native 1080p, with multilingual voice generation, precise lip-sync, cinematic camera control and shot transitions.
On April 13, 2026 Shengshu extended the same foundation with Vidu Q3 Reference-to-Video, which combines subjects, environments, costumes, props and visual styles from multiple references in one workflow, adds multi-shot composition, background music and sound effect generation, and six categories of cinematic effects such as particle systems, fluid simulation and lighting. The company reported that Q3 ranked first on the Artificial Analysis video benchmark at launch, and in the same announcement it closed a RMB 2 billion Series B led by Alibaba Cloud.
On Aura AI, Vidu Q3 runs in text-to-video and image-to-video mode with 5 second clips at 360p, 540p, 720p or 1080p in five aspect ratios, with native audio and no watermark. Reference work is handled by Vidu Q2 Reference in the reference-to-video tool, available on every plan. Both sit next to Kling O3, Seedance 2.5 and Pixverse 5 in the same studio.
Six things that set Shengshu's Q3 generation apart
Vidu Q3 integrates audio at the model level rather than adding it afterwards. Dialogue, music and effects come out of the same generation as the footage, already timed to the action.
Write a line in the language you need and the character speaks it with matching mouth movements. Useful for localized ads and short character moments without a separate dubbing step.
Q3 understands cinematic camera language and can cut between shots inside one clip. Describe a push-in, a pan or a transition and the model stages it rather than holding a single static frame.
With Vidu Q2 Reference on Aura AI, upload images of a character, a product or a location and keep them consistent across the clip. Starter allows 2 references, Pro 3 and Premium 5.
Shengshu's Q3 Reference-to-Video release covers six effect categories: particles, fluid simulation, dynamic motion, camera movement, transitions and lighting.
Four resolution tiers let you test a prompt at 2 credits (360p) or 3 credits (540p), then render the keeper at 720p for 4 credits or 1080p for 6.
Four steps from prompt to a finished clip with sound
Vidu Q3: 4 credits base for 5 seconds, multiplied by the resolution factor. Vidu Q2 Reference: 4 credits base for up to 5 seconds plus 0.5 per extra second, multiplied by its own factors (540p x0.7, 720p x1, 1080p x1.3)
| Model, duration | 360p (x0.5) | 540p (x0.7) | 720p (x1) | 1080p (x1.5 / x1.3) |
|---|---|---|---|---|
| Vidu Q3, 5 seconds | 2 credits | 3 credits | 4 credits | 6 credits |
| Vidu Q2 Reference, 3 seconds | Not offered | 3 credits | 4 credits | 6 credits |
| Vidu Q2 Reference, 5 seconds | Not offered | 3 credits | 4 credits | 6 credits |
| Vidu Q2 Reference, 8 seconds | Not offered | 4 credits | 6 credits | 8 credits |
How Shengshu's model compares to other short-clip options on Aura AI
| Feature | Vidu Q3 | Kling O3 | Seedance 2.5 | Pixverse 5 |
|---|---|---|---|---|
| Max resolution on Aura AI | 1080p | 1080p (4K via Kling O3 4K, Premium) | 1080p | 1080p |
| Duration | 5 s | 3 to 15 s | 4 to 15 s | 5 or 8 s |
| Native audio | ✓ Single pass, lip-sync | ✓ Yes | ✓ Single pass, 10+ languages | No |
| Reference to video on Aura AI | ✓ Vidu Q2 Reference, up to 5 images | Image + text | Image + text | Image + text |
| Resolution tiers | 4 (360p to 1080p) | Fixed | 3 (480p to 1080p) | 4 (360p to 1080p) |
| Aspect ratios | 5 | 3 | 6 | 5 |
| Credits, 5 s at 720p | 4 | From 3 | 7 | 1 |
| Plan | Starter+ | All plans | Starter+ | All plans |
Where native audio, lip-sync and reference consistency pay off
One prompt, several languages. Vidu Q3's multilingual voice generation with lip-sync lets you render the same 5 second spot for different markets without re-shooting or dubbing.
Upload two to five photos of a mascot, a presenter or a product in the reference-to-video tool and Vidu Q2 Reference keeps the look stable shot after shot.
Camera control plus shot transitions in one clip. A wide shot cutting to a close-up with the sound carrying across the cut reads like an edited scene rather than a single generated frame.
At 2 credits per 360p clip, Vidu Q3 is a cheap way to iterate on wording, camera moves and sound cues. Once the prompt works, re-render the same text at 1080p for 6 credits.
Common questions about Vidu Q3 on Aura AI
Video and audio in one pass, up to 1080p, no watermark. Vidu Q3 from the Starter plan, Vidu Q2 Reference on every plan, worldwide.