Shengshu Technology ยท Unveiled January 30, 2026

Vidu Q3 AI Video Generator

Shengshu's Vidu Q3 generates picture and sound in one pass, with multilingual lip-sync and camera control. 5 second clips from 360p to 1080p, five aspect ratios, no watermark. Available from the Starter plan on Aura AI, with Vidu Q2 Reference for reference-to-video on every plan.

1080p
Max Resolution
Audio
Single Pass
2
Credits From (360p)
5
Reference Images (Premium)

Vidu Q3 on Aura AI

MakerShengshu Technology
ModesText to video, image to video
Duration5 seconds
Resolution360p, 540p, 720p, 1080p
Aspect ratios16:9, 9:16, 4:3, 3:4, 1:1
AudioNative, with lip-sync
CreditsFrom 2 per clip
PlanStarter, Pro, Premium

What is Vidu Q3?

Vidu Q3 is the flagship video model of Shengshu Technology, the Beijing company that grew out of Tsinghua University's TSAIL lab and runs the Vidu platform. Shengshu unveiled Q3 on January 30, 2026 during its Global Creativity Week in Singapore, presenting it as the first long-form video model to deliver native audio and video in a single output. Upstream it generates up to 16 seconds of synchronized sound and picture at native 1080p, with multilingual voice generation, precise lip-sync, cinematic camera control and shot transitions.

On April 13, 2026 Shengshu extended the same foundation with Vidu Q3 Reference-to-Video, which combines subjects, environments, costumes, props and visual styles from multiple references in one workflow, adds multi-shot composition, background music and sound effect generation, and six categories of cinematic effects such as particle systems, fluid simulation and lighting. The company reported that Q3 ranked first on the Artificial Analysis video benchmark at launch, and in the same announcement it closed a RMB 2 billion Series B led by Alibaba Cloud.

On Aura AI, Vidu Q3 runs in text-to-video and image-to-video mode with 5 second clips at 360p, 540p, 720p or 1080p in five aspect ratios, with native audio and no watermark. Reference work is handled by Vidu Q2 Reference in the reference-to-video tool, available on every plan. Both sit next to Kling O3, Seedance 2.5 and Pixverse 5 in the same studio.

Vidu Q3 Capabilities

Six things that set Shengshu's Q3 generation apart

Sound and Vision Together

Vidu Q3 integrates audio at the model level rather than adding it afterwards. Dialogue, music and effects come out of the same generation as the footage, already timed to the action.

Multilingual Voices with Lip-Sync

Write a line in the language you need and the character speaks it with matching mouth movements. Useful for localized ads and short character moments without a separate dubbing step.

Camera Control and Shot Transitions

Q3 understands cinematic camera language and can cut between shots inside one clip. Describe a push-in, a pan or a transition and the model stages it rather than holding a single static frame.

Reference to Video

With Vidu Q2 Reference on Aura AI, upload images of a character, a product or a location and keep them consistent across the clip. Starter allows 2 references, Pro 3 and Premium 5.

Cinematic Effects Upstream

Shengshu's Q3 Reference-to-Video release covers six effect categories: particles, fluid simulation, dynamic motion, camera movement, transitions and lighting.

Low Cost Drafts at 360p

Four resolution tiers let you test a prompt at 2 credits (360p) or 3 credits (540p), then render the keeper at 720p for 4 credits or 1080p for 6.

How to Use Vidu Q3 on Aura AI

Four steps from prompt to a finished clip with sound

  1. Open the Text to Video tool, or Image to Video if you want to animate a still. For consistent characters from several photos, use Reference to Video with Vidu Q2 Reference instead. Sign in with your Aura AI account.
  2. Pick Vidu Q3 in the model selector. Choose a resolution (360p, 540p, 720p or 1080p) and one of the five aspect ratios. Clips are 5 seconds, and the credit cost updates live as you change the resolution.
  3. Write the prompt. Describe the subject, the camera move and any transition, put dialogue in quotes and name the music and ambience.
  4. Generate, preview and download the MP4 without watermark. For longer sequences, run a second generation from the last frame and join the clips in the AI video editor.

Vidu Q3 Pricing in Credits

Vidu Q3: 4 credits base for 5 seconds, multiplied by the resolution factor. Vidu Q2 Reference: 4 credits base for up to 5 seconds plus 0.5 per extra second, multiplied by its own factors (540p x0.7, 720p x1, 1080p x1.3)

Model, duration 360p (x0.5) 540p (x0.7) 720p (x1) 1080p (x1.5 / x1.3)
Vidu Q3, 5 seconds2 credits3 credits4 credits6 credits
Vidu Q2 Reference, 3 secondsNot offered3 credits4 credits6 credits
Vidu Q2 Reference, 5 secondsNot offered3 credits4 credits6 credits
Vidu Q2 Reference, 8 secondsNot offered4 credits6 credits8 credits
Vidu Q3 requires the Starter plan ($14/mo) or higher. Vidu Q2 Reference is available on every plan (Starter, Pro, Premium), with 2, 3 or 5 reference images respectively. Pro ($59/mo) and Premium ($99/mo) include larger monthly credit allowances. Fractions are rounded up to the next whole credit, and only successful generations are charged.

Vidu Q3 vs Kling O3 vs Seedance 2.5 vs Pixverse 5

How Shengshu's model compares to other short-clip options on Aura AI

Feature Vidu Q3 Kling O3 Seedance 2.5 Pixverse 5
Max resolution on Aura AI 1080p 1080p (4K via Kling O3 4K, Premium) 1080p 1080p
Duration 5 s 3 to 15 s 4 to 15 s 5 or 8 s
Native audio Single pass, lip-sync Yes Single pass, 10+ languages No
Reference to video on Aura AI Vidu Q2 Reference, up to 5 images Image + text Image + text Image + text
Resolution tiers 4 (360p to 1080p) Fixed 3 (480p to 1080p) 4 (360p to 1080p)
Aspect ratios 5 3 6 5
Credits, 5 s at 720p 4 From 3 7 1
Plan Starter+ All plans Starter+ All plans

What to Make with Vidu Q3

Where native audio, lip-sync and reference consistency pay off

Localized Spoken Ads

One prompt, several languages. Vidu Q3's multilingual voice generation with lip-sync lets you render the same 5 second spot for different markets without re-shooting or dubbing.

Example Prompt: "Young woman holding a skincare bottle in a bright bathroom, smiles at camera and says in Spanish 'Piel nueva en siete dias', soft pop track, 9:16, 1080p"

Consistent Characters from References

Upload two to five photos of a mascot, a presenter or a product in the reference-to-video tool and Vidu Q2 Reference keeps the look stable shot after shot.

Example Prompt: "The character from the references walks into a sunlit cafe, orders a coffee and sits by the window, warm tones, 16:9, 8 seconds"

Two-Shot Micro Scenes

Camera control plus shot transitions in one clip. A wide shot cutting to a close-up with the sound carrying across the cut reads like an edited scene rather than a single generated frame.

Example Prompt: "Wide shot of a surfer paddling out at sunrise, cut to close-up of her face as a wave rises behind, ocean roar, gulls, 16:9, 720p"

Rapid Prompt Testing

At 2 credits per 360p clip, Vidu Q3 is a cheap way to iterate on wording, camera moves and sound cues. Once the prompt works, re-render the same text at 1080p for 6 credits.

Example Prompt: "Drone shot over terraced rice fields at golden hour, slow forward glide, gentle wind and distant birds, 4:3, 360p"

Vidu Q3 FAQ

Common questions about Vidu Q3 on Aura AI

What is Vidu Q3? +
Vidu Q3 is the video generation model from Shengshu Technology, the Beijing company behind the Vidu platform, unveiled on January 30, 2026. It generates picture and sound in a single pass, with multilingual voices, lip-sync, camera control and shot transitions, and upstream produces clips up to 16 seconds at 1080p. On Aura AI it runs in text-to-video and image-to-video mode with 5 second clips at 360p, 540p, 720p or 1080p in five aspect ratios.
How many credits does Vidu Q3 cost on Aura AI? +
A 5 second Vidu Q3 clip costs 4 credits at 720p, 6 credits at 1080p, 3 credits at 540p and 2 credits at 360p. Vidu Q2 Reference starts at 4 credits for a 3 to 5 second clip at 720p plus 0.5 credits per extra second, with 540p at 0.7x and 1080p at 1.3x. Fractions are rounded up and only successful generations are charged.
Which plan do I need to use Vidu Q3? +
Vidu Q3 is available on the Starter ($14/mo), Pro ($59/mo) and Premium ($99/mo) plans. Vidu Q2 Reference in the reference-to-video tool is included on every plan. Each plan includes a monthly credit allowance and works on web, iOS, Android and Windows.
What is Vidu Q2 Reference and how many reference images can I use? +
Vidu Q2 Reference is Shengshu's reference-to-video model on Aura AI. You upload images of a character, a product or a setting and the model keeps them consistent in a 3 to 8 second clip at 540p, 720p or 1080p, in 16:9, 9:16 or 1:1, without audio. The reference limit depends on your plan: 2 images on Starter, 3 on Pro and 5 on Premium.
Does Vidu Q3 generate audio and lip-sync? +
Yes. Vidu Q3 was built to generate sound and vision together at the model level, so dialogue, background music and sound effects arrive synchronized with the footage, with multilingual voice generation and precise lip-sync. Put the spoken line in quotes and describe the music and ambience in your prompt.

Generate Vidu Q3 Videos Online

Video and audio in one pass, up to 1080p, no watermark. Vidu Q3 from the Starter plan, Vidu Q2 Reference on every plan, worldwide.