Black Forest Labs' first video model generates footage with dialogue, sound effects and ambience in one pass. 5 to 20 second clips at 720p or 1080p, six aspect ratios, no watermark. Available from the Starter plan on Aura AI.
FLUX 3 Video is the video side of FLUX 3, the model Black Forest Labs announced on July 23, 2026. Until now the Freiburg-based lab was known for image models, from the original FLUX.1 to FLUX.2. FLUX 3 changes the scope: a single set of weights trained jointly on images, video and audio, using a method BFL calls Self Flow, and extendable to predicting robot actions. Video opened first to early-access applicants and became generally available through the BFL API and selected partners on August 4, 2026.
Upstream, FLUX 3 Video generates clips up to 20 seconds at 720p or 1080p with native audio: dialogue, sound effects and ambient sound, in multiple languages, with lip-sync. It supports text-to-video and image-to-video, including keyframe specification and continuation of existing footage, and it can switch between scenes and camera angles inside a single generation. According to BFL's internal evaluation, human raters preferred FLUX 3 over existing models for text-to-video and it tied Seedance 2.0 on image-to-video. Those are vendor figures. An open-weight FLUX 3 Dev variant is on the roadmap without a fixed date.
On Aura AI, FLUX 3 Video runs in text-to-video and image-to-video mode with any duration from 5 to 20 seconds, at 720p or 1080p, in six aspect ratios from 21:9 to 9:16. Output is delivered without watermark and ready for commercial use. It sits next to Seedance 2.5, Veo 3.1 and Sora 2 Pro in the same studio, so you can compare takes from different models without switching tools.
Six things Black Forest Labs' first video model does well
Audio is part of the model, not a post-process. A prompt with a spoken line returns the voice lip-synced to the character, layered over sound effects and room tone, in the language you wrote it in.
FLUX 3 Video generates any length from 5 to 20 seconds on Aura AI, one second at a time. That covers a full Story, a product demo or a short scene without stitching clips together.
The model can cut between scenes and camera angles within a single generation. Describe two or three shots in the prompt and it stages them in order with consistent characters and lighting.
Start from a still to lock the look of a product, a character or a location. Upstream, FLUX 3 supports keyframe specification and continuation of existing footage on top of standard image-to-video.
FLUX 3 is a single multimodal model rather than a separate video pipeline.
21:9 ultrawide, 16:9, 4:3, 1:1, 3:4 and 9:16 vertical are all native on Aura AI, so the same prompt can be rendered for cinema, web, feed and Stories without cropping.
Four steps from prompt to a finished clip with sound
Cost = 6 credits base for 5 seconds, plus 0.5 credits per extra second, multiplied by the resolution factor
| Duration | 720p (x0.8) | 1080p (x1) |
|---|---|---|
| 5 seconds | 5 credits | 6 credits |
| 8 seconds | 6 credits | 8 credits |
| 10 seconds | 7 credits | 9 credits |
| 15 seconds | 9 credits | 11 credits |
| 20 seconds | 11 credits | 14 credits |
How BFL's first video model compares to the other flagship audio models on Aura AI
| Feature | FLUX 3 Video | Seedance 2.5 | Veo 3.1 | Sora 2 Pro |
|---|---|---|---|---|
| Max resolution on Aura AI | 1080p | 1080p | 4K | 1080p |
| Duration | 5 to 20 s | 4 to 15 s | 4, 6 or 8 s | 4, 8, 12, 16 or 20 s |
| Native audio | ✓ Dialogue, effects, lip-sync | ✓ Single pass, 10+ languages | ✓ Yes | ✓ Yes |
| Scene switches in one clip | ✓ Yes | ✓ Multi-shot | No | No |
| Multimodal references upstream | Image, keyframes, continuation | ✓ Up to 50 | Image + text | Image + text |
| Aspect ratios | 6 | 6 | 2 | 2 |
| Credits, 8 s at 1080p | 8 | 15 | 11 | 11 |
| Plan | Starter+ | Starter+ | Starter+ | Starter+ |
Where 20 second clips, dialogue and six ratios pay off
A 15 to 20 second walkthrough of a product with a spoken script and matching sound effects. Start from a packshot in image-to-video to keep the exact design.
Two characters, a written exchange, lip-synced in the language of your choice. FLUX 3 keeps faces and voices consistent across a scene switch, so a short exchange fits in one generation.
Native 9:16 with the soundtrack already mixed. Use the full 20 seconds for a Story-length ad, or 5 to 8 seconds for a punchy feed clip.
Pre-visualize a scene as a wide shot, an insert and a reaction in one clip. FLUX 3 stages the switches itself, which saves generating and stitching three separate takes.
Common questions about FLUX 3 Video on Aura AI
Black Forest Labs' first video model with native audio, up to 20 seconds at 1080p, no watermark. Available from the Starter plan, worldwide.