Alibaba's third-generation image model reads prompts up to 4,500 tokens and renders dense layouts, 10-pixel text and 12 languages. 2 credits to generate, 2 credits to edit, five aspect ratios, no watermark, on every Aura AI plan.
Qwen Image 3, officially Qwen-Image-3.0, is the third generation of the image model from Alibaba's Qwen team. It launched as an invite-only cloud preview on July 21, 2026 and opened to every user of the Qwen platform on August 5, 2026. The Qwen line started in August 2025 with Qwen-Image, a 20 billion parameter MMDiT model released with open weights under Apache 2.0 that became known for rendering text, especially Chinese, better than almost anything else at the time. Qwen-Image-Edit, a 20B editing model with semantic edits and pixel-level text rewriting, followed the same month, and Qwen-Image-2.0 arrived later with its own technical report and prompts of around 1,000 tokens.
Version 3.0 is built around one idea: images that work as a tool, not just a picture. It accepts prompts of up to 4,500 tokens, which is enough to describe a full newspaper page, a multi-panel storyboard or a dense infographic grid and get it in a single generation. Alibaba says it renders text as small as 10 pixels, reproduces textures approaching photographic quality, supports 12 languages natively, can draw web-page style interfaces and can pull live internet data into an image. Unlike its predecessors, 3.0 shipped without open weights, a parameter count, a license or a benchmark table, so it is a hosted-only model for now.
On Aura AI, Qwen Image 3 runs in text-to-image mode for 2 credits per image and, as Qwen Image 3 Edit, in the image editor for 2 credits per edit, in five aspect ratios, on every plan. It sits next to Hunyuan Image 3, the other Chinese text-rendering specialist, FLUX 2 and Nano Banana, so you can run the same layout prompt through several models and keep the best result, all without watermark.
Six capabilities that separate 3.0 from earlier Qwen-Image releases
Roughly four times the prompt budget of Qwen-Image-2.0. Describe every panel, every caption and every color in one message and the model composes the whole thing in a single pass.
Alibaba claims legible text down to 10 pixels tall. Fine print, table cells, axis labels and footnotes stay readable instead of dissolving into squiggles.
Newspapers, storyboards, comparison charts, infographic grids and app or web-page interfaces are first-class targets, not edge cases. The model keeps columns aligned and elements where you put them.
Text rendering covers 12 languages, building on the Chinese and English strength the first Qwen-Image was known for. Mixed-language signage and bilingual packaging come out coherent.
Qwen Image 3 Edit takes an existing image and a plain-language instruction. The Qwen editing line has supported semantic changes and in-place text rewriting since Qwen-Image-Edit in 2025.
Alibaba reports textures approaching photographic quality in 3.0, which closes the gap between the model's graphic-design strengths and the photoreal output of FLUX 2 or Nano Banana.
Four steps from a long prompt to a finished layout
Flat cost per image, the same in every aspect ratio
| Variant | Mode | Credits per image | Best for |
|---|---|---|---|
| Qwen Image 3 | Text to image | 2 credits | Infographics, posters, storyboards, UI mockups, multilingual text |
| Qwen Image 3 Edit | Image edit | 2 credits | Text fixes, panel swaps, semantic edits on existing images |
How Alibaba's layout specialist compares to its rivals on Aura AI
| Feature | Qwen Image 3 | Hunyuan Image 3 | FLUX 2 | Nano Banana |
|---|---|---|---|---|
| Maker | Alibaba | Tencent | Black Forest Labs | |
| Text-to-image credits | 2 | 1 | 2 (Turbo 1, Max 3) | 3 (Pro or 2) |
| Edit variant on Aura AI | ✓ 2 credits | No | ✓ 2 credits | ✓ 2 or 3 credits |
| Prompt length upstream | 4,500 tokens | 1,000+ characters | Standard | Standard |
| Text rendering | 12 languages, 10 px text | Chinese and English, strong | Good | Good |
| Open weights | No (1.0 was Apache 2.0) | ✓ Community License | Dev checkpoint only | No |
| Aspect ratios on Aura AI | 5 | 5 | 5 | 5 |
| Plan | All plans | All plans | All plans | All plans |
Where long prompts and tiny readable text pay off
Multi-section explainers with headings, numbered callouts and small labels, generated in one pass. Write the content as a list and the model lays it out.
Landing pages, dashboards and app screens with real-looking navigation, buttons and copy. Useful for pitching a concept before any design work starts.
Bilingual labels, storefront signs and menus where both scripts have to be correct. 12-language rendering handles mixed Latin and CJK text on the same surface.
Multi-panel pages with captions and dialogue in each frame, described panel by panel within the 4,500-token budget.
Common questions about Qwen Image 3 on Aura AI
Long prompts, dense layouts and readable text in 12 languages. 2 credits to generate, 2 to edit, no watermark. Available on every plan, worldwide.