Qwen-Image-2.1-Turbo: Alibaba's open image model now draws in 8 steps
Alibaba's Qwen team released Qwen-Image-2.1-Turbo on October 9, a faster checkpoint of the open image model it shipped in September. It draws and edits in 8 steps instead of 40, and keeps what made 2.1 stand out: transparent images, edits from up to 10 reference pictures, and 2K output. The catch is the license. Here's what's new, where to run it and how to prompt it.
What happened
On October 9, Alibaba's Qwen team added Qwen-Image-2.1-Turbo to Hugging Face, and the same day hosted Pro and Turbo versions went live in Alibaba Cloud's Model Studio API. Turbo is an accelerated checkpoint of Qwen-Image-2.1, released on September 20: the same 7-billion-parameter model that both generates and edits images, running 8 denoising steps where the base model uses 40. It went straight to the top of Hugging Face's trending list.
Qwen publishes no speed figure and no quality comparison between Turbo and the base model. Fewer steps should mean faster pictures on the same hardware, but that's arithmetic, not a measured claim.
Qwen-Image-2.1-Turbo released!
What 2.1 can do
- Transparent images. It makes PNGs with a real transparent background, edits individual transparent layers, including text inside one, and cuts a subject out of a photo as its own layer. Qwen used to ship that as a separate model; now it's built in.
- Up to 10 reference images. Qwen's examples build a group photo from six portraits, an outfit from five items and a furnished room from ten product shots.
- Point at what to change. Mark the area for a local edit by circling it in color, painting on it, or passing a separate mask.
- Faces, products and lettering. Qwen claims better identity keeping for people and products, and better typography, portrait lighting and fine detail.
- 2K by default. 2048×2048 for a square, with presets up to 2752×1536 for 16:9.
The license changed
Earlier Qwen-Image releases used Apache 2.0, which allows commercial use. Qwen-Image-2.1 and Turbo use the Qwen Research License: non-commercial use only, with a separate license needed for anything commercial, and "Built with Qwen" required on models made from it. In the Reddit launch thread, the change was the main complaint. If you want the images for paid work, the hosted API is the route Alibaba sells.
Where to run it
- ComfyUI. Supported natively since 2.1's release. Comfy-Org's repository already carries Turbo files, including a 14 GB BF16 model, a 7 GB int8 version and a Turbo LoRA.
- Your own code. Weights are on Hugging Face and ModelScope; Turbo needs a recent build of Diffusers. The full Turbo download is about 32 GB, most of it the text encoder. Qwen doesn't state a VRAM requirement and suggests CPU offload for smaller GPUs.
- Online. Qwen's Hugging Face Space runs Qwen-Image-2.1; we couldn't confirm it serves Turbo.
- API. On Alibaba Cloud Model Studio, Turbo costs 0.1 yuan per image, about 1.4 US cents, and Pro 0.25 yuan, at list price.
What it means for your prompts
- Write long. Qwen's own example prompts are dense paragraphs, the Turbo poster example runs to several hundred words, and any lettering goes in quotation marks. For short ideas, Qwen offers official prompt-rewriting models that expand them into a long English prompt and suggest an aspect ratio.
- Use the transparency phrase. For a transparent image, Qwen's template wraps your description: "This is an RGBA image with transparency" before it, "The image has alpha channel and the background is transparent" after it.
- Skip the negative prompt on Turbo. Turbo runs at a guidance scale of 1, and users in the launch thread point out that negative prompts stop having any effect there. Say what you want instead of what you don't.
- Running out of memory at 2K? The same thread says the image decode at the end is the usual culprit; turn on tiled decoding in ComfyUI.
Our AI image prompts are written in that long, descriptive style and work as a starting point.
Try it
This is an RGBA image with transparency. A [product, e.g. ceramic coffee mug in matte sage green] photographed from a slight three-quarter angle, centered, filling about two thirds of the frame. Soft studio light from the upper left, a gentle highlight on the rim, crisp edges, true-to-life color and texture. No shadow, no props, no surface underneath. The image has alpha channel and the background is transparent.
A vertical 3:4 poster for [event, e.g. a neighborhood jazz night]. At the top, the title "[TITLE, 3 words max]" in large condensed sans-serif capitals, [color] on a deep [color] background. Below it, a stylized illustration of [main image, e.g. a saxophone player in silhouette] in flat shapes with a subtle paper grain. At the bottom, two lines of small text, spelled exactly: "[date and time]" and "[place]". Generous margins, strong contrast, every letter sharp and correctly spelled, no other text anywhere.
Use the [number] images I've attached: [image 1: a person], [image 2: a jacket], [image 3: a pair of shoes], [image 4: a street]. Create one photo of the person from image 1 wearing the jacket from image 2 and the shoes from image 3, walking along the street from image 4 at golden hour. Keep the person's face, hair and build exactly as in image 1, and the jacket's color, logo and stitching as in image 2. Full-body shot at eye level, natural motion, realistic light and shadows, 2:3.
Sources
- Qwen, Qwen-Image-2.1-Turbo model card, October 9, 2026: what Turbo is, 8 steps, resolutions, requirements
- Qwen, Qwen-Image-2.1 on GitHub: release dates, architecture, transparency template, prompt rewriters, where it runs
- Qwen, Qwen-Image-2.1 announcement, September 20, 2026: new capabilities
- Qwen, Qwen Research License: non-commercial terms
- Comfy-Org, Qwen-Image-2.1 files for ComfyUI: Turbo checkpoints and LoRA
- Alibaba Cloud, Qwen-Image-2.1-Turbo in Model Studio: API price and limits
- Reddit, r/LocalLLaMA launch thread, October 9, 2026: license reaction, user tips