ByteDance's newest video model generates picture and sound in one pass. 4 to 15 second clips, 480p to 1080p, six aspect ratios, no watermark. Available from the Starter plan on Aura AI.
Seedance 2.5 is the July 2026 release of ByteDance's Seedance video family, announced on June 23 and made available on July 31, 2026. It follows Seedance 2.0, the model that introduced native audio and dance mode earlier in the year, and it pushes the same idea further: one generation pass produces the footage, the dialogue, the music and the sound effects together, already in sync.
The headline changes are about control rather than raw resolution. Seedance 2.5 accepts up to 50 multimodal references in a single job upstream (30 images, 10 video clips and 10 audio tracks), each of which can be assigned a role such as character identity, product appearance, camera style or voice. It adds director-style multi-shot camera direction, region-level editing of existing clips, extension before or after a clip, and ByteDance reports roughly 20 percent better prompt adherence than 2.0. Independent benchmarks are still pending, so treat that figure as a vendor claim.
On Aura AI, Seedance 2.5 runs in text-to-video and image-to-video mode with clips from 4 to 15 seconds at 480p, 720p or 1080p, in six aspect ratios from 21:9 ultrawide to 9:16 vertical. Output is delivered without watermark and is ready for commercial use. It sits alongside Seedance 2.0, Seedance 1.5, Kling 3.0, Veo 3.1 and 30 other models in the same studio, so you can switch models without changing tools.
Six capabilities that separate 2.5 from earlier Seedance releases
Dialogue, music and ambient sound are generated together with the picture. Lip-sync works in more than 10 languages, so a talking character or a sung line arrives already aligned to the mouth movements.
Seedance 2.5 accepts 30 images, 10 video clips and 10 audio tracks per generation upstream, each tagged with a role. A single audio reference can drive pacing, beat matching and lip-sync for the entire clip.
Describe a sequence of shots, camera moves and cuts in one prompt and the model stages them in order. Characters, lighting and style stay consistent from the establishing shot to the close-up.
Region-level editing changes a specific area of an existing clip while leaving the rest untouched. Swap a product, change a costume or fix a background without re-rolling the whole scene.
Add footage at the head or tail of a clip. The model preserves the characters, rhythm and visual style of the source, which is the practical way to stretch a 15 second clip into a longer sequence.
ByteDance reports about 20 percent better instruction following than Seedance 2.0. In practice that means fewer re-rolls for multi-subject scenes, specific camera language and precise timing of actions.
Four steps from prompt to a finished clip with sound
Cost = 6 credits base for 4 seconds, plus 0.35 credits per extra second, multiplied by the resolution factor
| Duration | 480p (x0.6) | 720p (x1) | 1080p (x2) |
|---|---|---|---|
| 4 seconds | 4 credits | 6 credits | 12 credits |
| 8 seconds | 5 credits | 8 credits | 15 credits |
| 10 seconds | 5 credits | 9 credits | 17 credits |
| 15 seconds | 6 credits | 10 credits | 20 credits |
How the new ByteDance release compares to its siblings on Aura AI
| Feature | Seedance 2.5 | Seedance 2.0 | Kling 3.0 | Veo 3.1 |
|---|---|---|---|---|
| Max resolution on Aura AI | 1080p | 4K | 1080p (4K via Kling V3 4K) | 4K |
| Duration | 4 to 15 s | 4 to 15 s | 3 to 15 s | 4, 6 or 8 s |
| Native audio | ✓ Single pass, 10+ languages | ✓ Stereo | ✓ Yes | ✓ Yes |
| Multimodal references | ✓ Up to 50 (upstream) | ✓ Up to 12 | Image + text | Image + text |
| Edit and extend clips | ✓ Region edit, extend | ✓ Yes | Extend | No |
| Aspect ratios | 6 | 6 | 3 | 2 |
| Credits, 8 s at 720p | 8 | 7 | 5 | 9 |
| Plan | Starter+ | Starter+ | All plans | Starter+ |
Where single-pass audio and multi-shot control pay off
Two characters, a written exchange, lip-synced in the language of your choice. Seedance 2.5 keeps voices and faces consistent across shots, which makes short narrative scenes practical in one generation.
9:16 product clips with voice-over and music already mixed. Use image-to-video with a product photo to lock the exact look, then let the model add motion and a spoken line.
The Seedance line was built for choreography. Describe the track and the moves and 2.5 will match cuts and body motion to the beat, with sung lyrics lip-synced when you include them.
Pre-visualize a scene as an establishing shot, an action beat and a reaction shot in a single clip. Director-style prompting keeps the same characters and lighting through every cut.
Common questions about Seedance 2.5 on Aura AI
Video and audio in one pass, up to 1080p, no watermark. Available from the Starter plan, worldwide.