Google's fast Gemini-family video model. Text to video, image to video and reference to video with native audio, 3 to 10 second clips at 720p, no watermark. From 3 credits on every Aura AI plan.
Gemini Omni Flash is Google's video generation and editing model built directly on the Gemini family rather than on the Veo line. Google previewed it in May 2026 and shipped the production release, Gemini Omni 1.1 Flash, on August 27, 2026 through the Gemini API, Google AI Studio, Flow and the Gemini app. Google describes it as a model for fast, conversational video generation: generate a clip, then refine it by describing the change in plain language.
Upstream, Gemini Omni Flash outputs 3 to 10 second clips at 24 fps in 360p, 720p, 1080p or 4K, extends a scene in 10 second steps up to 40 seconds, accepts a first and a last frame to define a shot, and takes short video clips as references. Artificial Analysis placed it first for text-to-video and in the top two for image-to-video at the end of July 2026, ahead of MiniMax H3 in both. The Flash name follows the Gemini convention for the tier tuned for speed and cost.
On Aura AI, Gemini Omni Flash runs in three modes. Text-to-video and image-to-video produce 3 to 10 second clips at 720p in 16:9 or 9:16 with native audio, and reference-to-video locks a character or product from 2 to 5 reference images depending on your plan. Every clip costs 3 credits plus 0.3 credits per second beyond 4 seconds, with no resolution multiplier, on every plan. It shares the selector with Veo 3.1 Lite, Grok Imagine and Seedance 2.5.
Six reasons to route a prompt through Google's Gemini video model
Ambience, sound effects and dialogue are generated together with the picture and stay aligned with the motion. Put spoken lines in quotes and describe the room tone, and the clip arrives ready to post.
A 4 second clip with audio is 3 credits and a 10 second clip is 5. That makes it the right tool for iterating on a prompt before committing to a Veo 3.1 render.
Attach 2 to 5 images of a person, mascot or product and the model keeps them consistent across the clip. Reference-to-video costs the same as text-to-video, so character-locked shots do not carry a premium.
Veo clips come in fixed 4, 6 or 8 second lengths. Gemini Omni Flash accepts any whole number of seconds between 3 and 10, so a 7 second hook or a 9 second product loop is generated at exactly that length.
Google built Omni Flash around refining a video by talking to it: change the lighting, extend the scene, swap the ending. Aura AI exposes the generation modes today.
The Artificial Analysis text-to-video ranking had Gemini Omni Flash in first place at the end of July 2026, with image-to-video in the top two. For a model priced from 3 credits, that is unusual.
Four steps from prompt to a 720p clip with sound
Cost = 3 credits base for up to 4 seconds, plus 0.3 credits per extra second, rounded up. Single 720p tier, no resolution multiplier
| Duration | Text to video (720p) | Image to video (720p) | Reference to video (720p) |
|---|---|---|---|
| 3 seconds | 3 credits | 3 credits | 3 credits |
| 4 seconds | 3 credits | 3 credits | 3 credits |
| 5 seconds | 4 credits | 4 credits | 4 credits |
| 6 seconds | 4 credits | 4 credits | 4 credits |
| 8 seconds | 5 credits | 5 credits | 5 credits |
| 10 seconds | 5 credits | 5 credits | 5 credits |
How Google's Gemini video model compares to the other audio-capable models available on every plan
| Feature | Gemini Omni Flash | Veo 3.1 Lite | Veo 3 | Grok Video |
|---|---|---|---|---|
| Maker | xAI | |||
| Modes on Aura AI | Text, image, reference | Text, image | Text, image | Text, image, reference |
| Max resolution on Aura AI | 720p | 1080p | 1080p | 720p |
| Duration | 3 to 10 s, any length | 4, 6 or 8 s | 4, 6 or 8 s | 3 to 15 s |
| Native audio | ✓ Yes | ✓ Yes | ✓ Yes | ✓ Yes |
| Pricing model | 3 + 0.3 per extra second | Flat 2 | 6 + 1 per extra second | 3 + 0.3 per extra second |
| Credits, 8 s at 720p | 5 | 2 | 8 | 4 |
| Plan | All plans | All plans | All plans | All plans |
Where a fast, audio-capable model with flexible lengths pays off
A 5 to 7 second vertical opener with a spoken line and matching ambience. Gemini Omni Flash generates it at the exact length for 4 credits, audio included, so you can produce ten variants and test them.
Use reference-to-video with 2 to 5 photos of the same presenter or mascot and generate a series of scenes with a consistent face and outfit, priced like text-to-video.
Image-to-video turns a product shot or a landscape photo into a 720p clip with the right sound bed. 3 credits covers a 4 second loop.
Test framing, pacing and dialogue on Gemini Omni Flash at 3 to 5 credits per take, then send the final prompt to Veo 3.1 for a 1080p or 4K render.
Common questions about Gemini Omni Flash on Aura AI
Google's fast Gemini video model with native audio, 3 to 10 seconds at 720p, no watermark. From 3 credits on every plan, worldwide.