LTX 2.3 Review: 20-Second Audio Clips vs Wan 3.0 and Vidu Q3

LTX 2.3 is the least expensive way to get native 4K video with sound on Aura AI, and the only model in the catalog that runs a 20 second clip in one generation. Lightricks released it in March 2026 as an open-weight, 22 billion parameter model that produces the footage and a synchronized stereo soundtrack in a single pass, and on Aura AI it ships in two settings: LTX 2.3 for 6 to 10 second clips and LTX 2.3 Fast for 6 to 20 seconds, from 2 to 3 credits at 1080p on every plan. This review covers what the model does, what it costs, and how it compares with Alibaba's Wan 3.0 and Shengshu's Vidu Q3, the two other single-pass audio models released around it.

LTX 2.3 vs Wan 3.0 vs Vidu Q3 at a glance

All three generate picture and sound together. The table shows what each exposes on Aura AI; credits are per clip, rounded up.

Model Credits Resolution on Aura AI Duration Audio Plan
LTX 2.3 3 for 6 s + 0.3/s at 1080p (x1.5 at 1440p, x2.5 at 4K) 1080p, 1440p, 2160p 6, 8 or 10 s Stereo, single pass Every plan
LTX 2.3 Fast 2 for 6 s + 0.2/s at 1080p (same multipliers) 1080p, 1440p, 2160p 6 to 20 s Stereo, single pass Every plan
Wan 3.0 4 / 5 / 6 at 480p / 720p / 1080p 480p, 720p, 1080p 5 s Native Starter+
Wan 3.0 Prime 5 / 6 / 7 at 480p / 720p / 1080p 480p, 720p, 1080p 5 s Native, faster Starter+
Vidu Q3 2 / 3 / 4 / 6 at 360p / 540p / 720p / 1080p 360p to 1080p 5 s Native, multilingual lip-sync Starter+

Pricing tables and every setting are on the LTX 2.3 model page, the Wan 3.0 model page and the Vidu Q3 model page.

What LTX 2.3 is

LTX 2.3 is the March 2026 release of the LTX-2 line from Lightricks. It is a 22 billion parameter diffusion transformer that treats video and audio as one task: a single forward pass returns the footage and a synchronized stereo track, with dialogue, ambience and effects timed by construction rather than by a second tool. Lightricks positions it as the first open-weight model to combine native 4K output with joint audio generation, and publishes the checkpoints on Hugging Face under the LTX-2 Community License, so the same model can be run on-premises.

Compared with the original LTX-2 from January 2026, version 2.3 rebuilds the VAE for crisper textures and more natural movement, quadruples the size of the text connector with gated attention for better prompt adherence, adds explicit camera language (dolly, jib, focus shift) and cleans up audio artifacts. Third-party reviews describe the practical result as sharper detail across the whole frame, cleaner sound and stronger prompt following. Upstream the model generates at 24, 25, 48 or 50 frames per second, in landscape or native portrait, with text to video, image to video, audio to video and video to video modes.

On Aura AI it runs in text to video and image to video, in 16:9 or 9:16, at 1080p, 1440p or true 2160p, without watermark, on every plan. There is no GPU to rent and no installation.

20 second clips: how LTX 2.3 Fast works

The standard LTX 2.3 setting generates 6, 8 or 10 second clips. LTX 2.3 Fast trades a little quality for lower cost and longer durations: 6 to 20 seconds in 2 second steps, at 2 credits base plus 0.2 per extra second. A full 20 second clip at 1080p is 5 credits, at 1440p 8, at 4K 12. Native 9:16 means a complete vertical Story or Short arrives in one generation with its own sound, without cropping a landscape frame.

That length is the headline against the field. Wan 3.0 produces 30 second clips upstream and Vidu Q3 up to 16, but on Aura AI both are exposed as 5 second clips; Seedance 2.5 and Kling 3.0 stop at 15. Anything past 20 seconds still means chaining from the last frame and joining in the video editor, but LTX 2.3 Fast halves the number of joins.

4K with audio and what it costs

LTX 2.3 renders 2160p directly instead of upscaling a smaller frame. On Aura AI the 4K option costs 2.5 times the 1080p price: a 10 second 4K clip is 11 credits on LTX 2.3 and 7 on Fast, a 6 second 4K clip is 8 or 5. For context, the model page's own comparison puts an 8 second 1080p clip at 4 credits on LTX 2.3 (3 on Fast) against 5 on Kling 3.0 and 15 on Seedance 2.5, and lists no 4K at all on Seedance 2.5, standard Kling 3.0 or Veo 3.1 Lite.

Wan 3.0 and Vidu Q3 top out at 1080p on Aura AI. If the deliverable is a trade show wall, a lobby display or a stream background, LTX 2.3 is the only one of the three that produces the file at the resolution the screen needs.

LTX 2.3 vs Wan 3.0

Wan 3.0 comes from Alibaba's Tongyi Lab and opened in public beta on August 6, 2026. Its signature feature is Omni-Reference: up to 20 references in one job upstream, spanning images, video, audio, documents and webpage URLs, so the model can build a clip from a deck or a product page. Alibaba also reports more expressive faces and micro-expressions, steadier subject references and cleaner on-screen text than Wan 2.7. Unlike the openly released Wan 2.2, Wan 3.0 is a closed, API-only model with no published weights. Wan 3.0 Prime is the accelerated variant, one credit more per clip on Aura AI.

Against LTX 2.3 the trade is clear. Wan 3.0 offers five aspect ratios (16:9, 4:3, 1:1, 3:4, 9:16) where LTX offers two, and a 5 second 1080p clip is 6 credits against 2 on LTX 2.3 Fast for 6 seconds. LTX goes to 20 seconds and 4K; Wan stays at 5 seconds and 1080p. Wan 3.0 needs the Starter plan, LTX 2.3 is on every plan.

Pick

Wan 3.0 for a 5 second square or 4:3 product clip with a readable label and an expressive face. LTX 2.3 for anything longer, anything 4K, or anything you generate in volume.

LTX 2.3 vs Vidu Q3

Vidu Q3 is the flagship of Shengshu Technology, unveiled on January 30, 2026 as a long-form model that delivers native audio and video in a single output, with multilingual voice generation, precise lip-sync, cinematic camera control and shot transitions. Shengshu reported that it ranked first on the Artificial Analysis video benchmark at launch. On Aura AI it runs 5 second clips at 360p, 540p, 720p or 1080p in five aspect ratios, from 2 credits, and Vidu Q2 Reference handles reference-to-video on every plan.

Vidu Q3's strengths are the spoken line and the cheap test. A quoted sentence arrives lip-synced in the language you asked for, and a 360p draft costs 2 credits, the same as a 6 second LTX 2.3 Fast clip at 1080p. LTX 2.3 wins on length, on 4K and on price per second at full resolution: 20 seconds of 1080p for 5 credits against 5 seconds for 6. Both are built for audio; only Vidu documents multilingual lip-sync as a headline feature.

Pick

Vidu Q3 for a character delivering dialogue, for shot transitions inside one clip, and for iterating on a prompt at 360p. LTX 2.3 for long vertical clips, 4K backgrounds and high-volume B-roll.

Verdict

LTX 2.3 earns its place on three counts: the longest single generation on Aura AI, the cheapest native 4K, and open weights that mean the same model is available if you ever move a pipeline in-house. Its limits are two aspect ratios and, on the Fast setting, a small quality trade for length. Wan 3.0 is the better 5 second clip for framing variety and document-driven references; Vidu Q3 is the better 5 second clip for a spoken line. Because all three sit in the same studio, the sensible workflow is to draft on Vidu Q3 at 360p or LTX 2.3 Fast at 1080p, then render the keeper on the model whose strength matches the shot.

How to generate LTX 2.3 clips in Aura AI

Four steps from prompt to a 4K clip with sound.

  1. Open Text to Video, or Image to Video to animate a still, and sign in.
    LTX 2.3 and LTX 2.3 Fast appear in the model selector on every plan.
  2. Pick LTX 2.3 for 6, 8 or 10 seconds, or LTX 2.3 Fast for anything up to 20.
    Choose 1080p, 1440p or 2160p and either 16:9 or 9:16. The credit cost updates live.
  3. Write the prompt with the camera and the sound in it.
    Name the move (dolly in, jib up, rack focus), describe the subject and the environment, put a line of dialogue in quotes and name the ambience and music style.
  4. Generate, preview and download the MP4 without watermark.
    For sequences longer than 20 seconds, chain a second clip from the last frame and join them in the AI video editor.

Generate LTX 2.3 videos online

Native 4K with audio, up to 20 seconds with Fast, no watermark, from 2 credits. Included on every plan, worldwide.

Open Text to Video

Frequently Asked Questions

What is LTX 2.3?

LTX 2.3 is the March 2026 release of Lightricks' LTX-2 video model, a 22 billion parameter diffusion transformer that generates video and synchronized stereo audio in a single pass at native 4K. Its weights are published on Hugging Face under the LTX-2 Community License. On Aura AI it runs as LTX 2.3 (6, 8 or 10 second clips) and LTX 2.3 Fast (6 to 20 seconds) in text to video and image to video, at 1080p, 1440p or 2160p, on every plan.

How many credits does LTX 2.3 cost on Aura AI?

LTX 2.3 starts at 3 credits for a 6 second clip at 1080p plus 0.3 credits per extra second. LTX 2.3 Fast starts at 2 credits plus 0.2 per extra second. 1440p costs 1.5 times the 1080p price and 2160p costs 2.5 times. A 10 second 4K clip is 11 credits on LTX 2.3 or 7 on Fast, and a 20 second 1080p clip on Fast is 5 credits. Only successful generations are charged.

Can LTX 2.3 really generate 20 second clips?

Yes, through LTX 2.3 Fast, which runs from 6 to 20 seconds in 2 second steps with native audio, in 16:9 or 9:16, at 1080p, 1440p or 2160p. The standard LTX 2.3 setting offers 6, 8 or 10 seconds. No other model on Aura AI exposes 20 seconds in one generation; Seedance 2.5 and Kling 3.0 stop at 15, Wan 3.0 and Vidu Q3 at 5.

LTX 2.3 or Wan 3.0: which should I pick?

Pick LTX 2.3 for anything longer than 5 seconds, for 4K, or for the lowest credit cost: a 6 second 1080p clip is 2 credits on Fast against 6 for Wan 3.0. Pick Wan 3.0 for a 5 second clip in 4:3, 1:1 or 3:4, for readable on-screen text and expressive faces, or when its Omni-Reference inputs matter to you. Wan 3.0 needs the Starter plan; LTX 2.3 is on every plan.

LTX 2.3 or Vidu Q3: which should I pick?

Pick Vidu Q3 for a spoken line with precise multilingual lip-sync, for shot transitions, or for 2 credit prompt tests at 360p. Pick LTX 2.3 for 4K, for clips longer than 5 seconds and for the cheapest 1080p per second. Both generate audio with the picture; Vidu Q3 needs the Starter plan, LTX 2.3 is included on every plan.