Alibaba's Qwen-Image-3.0 Pushes Long-Context Image Generation to 4.5K Tokens
On July 21, 2026, Alibaba released Qwen-Image-3.0, the third generation of its image generation . The new version is built for very long text prompts and can accept an of up to about 4,500 tokens in a single pass. A single prompt can now describe a complex scene with multiple paragraphs of instructions and still produce a coherent image. The biggest change is the way the model handles text inside images. Earlier models often produced blurry or broken letters when asked to write sentences, formulas, or interface labels. Qwen-Image-3.0 can clean text directly, including mathematical formulas, geometric shapes, and step-by-step logic diagrams. It also generates complex user-interface mockups, where buttons, menus, and labels all need to be sharp and readable. Language support has also expanded. The model renders twelve languages natively, including English, Chinese, Japanese, Korean, Arabic, and Hindi. It ships with more than twenty built-in fonts, so the text it writes inside images looks consistent and professional. Localisation teams no longer need a separate tool to add translated captions or signs. The is part of a wider race among Chinese AI labs to build long-context image models. Rivals such as Tencent and ByteDance are also working on systems that can read long instructions and return finished visual assets. Alibaba is betting that a model which reads like a person and writes like a designer will win the next round of enterprise contracts in advertising, education, and software design.
/単語 · クリックで意味を確認
/確認クイズ 5 問
1. When did Alibaba release Qwen-Image-3.0?
2. Roughly how long is the single-pass text input that the new model can accept?
3. What kind of in-image content can the new model render cleanly, according to the article?
4. How many languages does the model support natively?
5. What does Alibaba believe the next round of enterprise contracts will reward?