Lightricks' open-weight video model generates native 4K footage and synchronized audio in one pass. 6 to 20 second clips at 1080p, 1440p or 2160p, no watermark, from 2 credits. LTX 2.3 and LTX 2.3 Fast are available on every Aura AI plan.
LTX 2.3 is the March 2026 release of the LTX-2 line from Lightricks. It is a 22 billion parameter diffusion transformer that treats video and audio as one generation task: a single forward pass returns the footage and a synchronized stereo soundtrack. Lightricks positions it as the first open-weight model to combine native 4K output with joint audio generation, and the weights are published on Hugging Face under the LTX-2 Community License, so it can also be run on-premises.
Compared with the original LTX-2 from January 2026, version 2.3 rebuilds the VAE for crisper textures and more natural movement, quadruples the size of the text connector with gated attention for better prompt adherence, adds motion and camera controls such as dolly, jib and focus shift, and cleans up audio artifacts. Upstream the model generates at 24, 25, 48 or 50 frames per second, in clips up to 20 seconds, in landscape and native portrait orientation, with text-to-video, image-to-video, audio-to-video and video-to-video modes.
On Aura AI, LTX 2.3 comes in two settings. LTX 2.3 generates 6, 8 or 10 second clips; LTX 2.3 Fast trades a little quality for lower cost and longer durations, up to 20 seconds. Both output 1080p, 1440p or true 2160p in 16:9 or 9:16 with native audio, without watermark, and both are included on every plan. That makes LTX 2.3 the least expensive way to get 4K on Aura AI, next to Seedance 2.5, Kling 3.0 and Veo 3.1 in the same studio.
Six reasons to pick LTX 2.3 for 4K clips with sound
LTX 2.3 renders 2160p directly instead of upscaling a smaller frame. On Aura AI you choose 1080p, 1440p or 2160p per generation, in landscape 16:9 or vertical 9:16, and get a 4K MP4 ready for large screens.
Dialogue, ambience and effects are generated together with the picture as one task, so timing matches by construction. No second tool, no manual sync.
The Fast setting on Aura AI starts at 2 credits and stretches to 20 second clips in 2 second steps.
Version 2.3 adds explicit camera language: dolly, jib and focus shift are understood from the prompt, and multi-subject layouts stay consistent across the whole clip.
A text connector four times larger than in LTX-2, with gated attention, reads longer prompts more precisely.
Lightricks publishes the LTX 2.3 checkpoints on Hugging Face under the LTX-2 Community License. On Aura AI you get the same model hosted, without a GPU, and the output is yours to use commercially.
Four steps from prompt to a 4K clip with sound
LTX 2.3: 3 credits base for 6 seconds plus 0.3 per extra second. Fast: 2 credits base plus 0.2 per extra second. Multiplied by the resolution factor
| Model, duration | 1080p (x1) | 1440p (x1.5) | 2160p / 4K (x2.5) |
|---|---|---|---|
| LTX 2.3, 6 seconds | 3 credits | 5 credits | 8 credits |
| LTX 2.3, 8 seconds | 4 credits | 6 credits | 9 credits |
| LTX 2.3, 10 seconds | 5 credits | 7 credits | 11 credits |
| LTX 2.3 Fast, 6 seconds | 2 credits | 3 credits | 5 credits |
| LTX 2.3 Fast, 10 seconds | 3 credits | 5 credits | 7 credits |
| LTX 2.3 Fast, 14 seconds | 4 credits | 6 credits | 9 credits |
| LTX 2.3 Fast, 20 seconds | 5 credits | 8 credits | 12 credits |
How Lightricks' open-weight model compares to other audio-capable models on Aura AI
| Feature | LTX 2.3 | Seedance 2.5 | Kling 3.0 | Veo 3.1 Lite |
|---|---|---|---|---|
| Max resolution on Aura AI | 2160p (4K) | 1080p | 1080p (4K via Kling V3 4K, Premium) | 1080p |
| Duration | 6 to 10 s (Fast: 6 to 20 s) | 4 to 15 s | 3 to 15 s | 4, 6 or 8 s |
| Native audio | ✓ Stereo, single pass | ✓ Single pass, 10+ languages | ✓ Yes | ✓ Yes |
| Open weights | ✓ Hugging Face | No | No | No |
| Aspect ratios | 2 | 6 | 3 | 2 |
| Credits, 8 s at 1080p | 4 (Fast 3) | 15 | 5 | 2 |
| Credits, 10 s at 4K | 11 (Fast 7) | No 4K | No 4K on this model | No 4K |
| Plan | All plans | Starter+ | All plans | All plans |
Where native 4K, long Fast clips and low credit cost pay off
Trade show walls, lobby displays and stream backgrounds need real 2160p. Generate a slow atmospheric shot at 4K with ambient sound, then loop it in the video editor.
LTX 2.3 Fast covers a full 20 second Story or Short in one generation at 5 credits for 1080p. Native 9:16 means no cropping of a landscape frame.
Joint audio generation makes a spoken line and its room tone arrive together. Good for explainer snippets, character intros and social ads with a single quoted sentence.
At 2 to 3 credits per 1080p clip, LTX 2.3 Fast is the setting to run many variations of an establishing shot and pick the best take without watching the budget.
Common questions about LTX 2.3 on Aura AI
Native 4K with audio, up to 20 seconds with Fast, no watermark. Included on every plan, worldwide.