Model reference · open weights

Qwen-Image-Edit-2509

Image Qwen Image edit 2 builds Open weights 463k dl/mo

Qwen-Image-Edit-2509 is an open-weight image model from Qwen. Qwen-Image-Edit-2509-GGUF (GGUF) weighs 13.1 GB; the smallest configuration that runs it is RTX 4090 24 GB.

Qwen-Image-Edit-2509 is a 20.4B parameter image-to-image model developed by Qwen for editing single or multiple images. It supports multi-image inputs, such as person and product combinations, and features native ControlNet support for depth, edge, and keypoint maps. The model is licensed under Apache 2.0 and operates in English and Chinese.

Summary of the Qwen/Qwen-Image-Edit-2509 model card, 2026-10-01 — the estimate below is for another build of the family

What it is

Released byQwen
TypeImage models
TaskImage edit
Parameters (lead)20.4B
Runs withdiffusers
Released2025-09-22
Popularity463k downloads / month
Weights13.1 GB (Qwen-Image-Edit-2509-GGUF (GGUF), file size)
LicenceOpen weights

What it runs on

Memory and cards for Qwen-Image-Edit-2509-GGUF (GGUF)

Weights 13.1 GB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GB … RTX 4060 Ti 16 GBdoes not fit
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

Builds

Sizes, precisions & builds

BuildParametersPrecisionWeightsSmallest card (1 image)
Qwen-Image-Edit-2509 ↗ 20.4BBF16 57.7 GBH100 80 GB
Qwen-Image-Edit-2509-GGUF (above) ↗
packaged by unsloth
16 builds: Q2_K 7.2 GB … F16 40.9 GB
the transformer alone — plus the text encoders and VAE
—GGUF 13.1 GBRTX 4090 24 GB

Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.

From the model card

What Qwen says about Qwen-Image-Edit-2509

Read the model card

💜 Qwen Chat&nbsp&nbsp | &nbsp&nbsp🤗 Hugging Face&nbsp&nbsp | &nbsp&nbsp🤖 ModelScope&nbsp&nbsp | &nbsp&nbsp 📑 Tech Report &nbsp&nbsp | &nbsp&nbsp 📑 Blog &nbsp&nbsp 🖥️ Demo&nbsp&nbsp | &nbsp&nbsp💬 WeChat (微信)&nbsp&nbsp | &nbsp&nbsp🫨 Discord&nbsp&nbsp| &nbsp&nbsp Github&nbsp&nbsp

Introduction

This September, we are pleased to introduce Qwen-Image-Edit-2509, the monthly iteration of Qwen-Image-Edit. To experience the latest model, please visit Qwen Chat and select the "Image Editing" feature. Compared with Qwen-Image-Edit released in August, the main improvements of Qwen-Image-Edit-2509 include:

  • Multi-image Editing Support: For multi-image inputs, Qwen-Image-Edit-2509 builds upon the Qwen-Image-Edit architecture and is further trained via image concatenation to enable multi-image editing. It supports various combinations such as "person + person," "person + product," and "person + scene." Optimal performance is currently achieved with 1 to 3 input images.
  • Enhanced Single-image Consistency: For single-image inputs, Qwen-Image-Edit-2509 significantly improves editing consistency, specifically in the following areas:
    • Improved Person Editing Consistency: Better preservation of facial identity, supporting various portrait styles and pose transformations;
    • Improved Product Editing Consistency: Better preservation of product identity, supporting product poster editing;
    • Improved Text Editing Consistency: In addition to modifying text content, it also supports editing text fonts, colors, and materials;
  • Native Support for ControlNet: Including depth maps, edge maps, keypoint maps, and more.

Quick Start

Install the latest version of diffusers

pip install git+https://github.com/huggingface/diffusers

The following contains a code snippet illustrating how to use Qwen-Image-Edit-2509:

import os
import torch
from PIL import Image
from diffusers import QwenImageEditPlusPipeline

pipeline = QwenImageEditPlusPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", torch_dtype=torch.bfloat16)
print("pipeline loaded")

pipeline.to('cuda')
pipeline.set_progress_bar_config(disable=None)
image1 = Image.open("input1.png")
image2 = Image.open("input2.png")
prompt = "The magician bear is on the left, the alchemist bear is on the right, facing each other in the central park square."
inputs = {
    "image": [image1, image2],
    "prompt": prompt,
    "generator": torch.manual_seed(0),
    "true_cfg_scale": 4.0,
    "negative_prompt": " ",
    "num_inference_steps": 40,
    "guidance_scale": 1.0,
    "num_images_per_prompt": 1,
}
with torch.inference_mode():
    output = pipeline(**inputs)
    output_image = output.images[0]
    output_image.save("output_image_edit_plus.png")
    print("image saved at", os.path.abspath("output_image_edit_plus.png"))

Showcase

The primary update in Qwen-Image-Edit-2509 is support for multi-image inputs.

Let’s first look at a "person + person" example:

Here is a "person + scene" example:

Below is a "person + object" example:

In fact, multi-image input also supports commonly used ControlNet keypoint maps—for example, changing a person’s pose:

Similarly, the following examples demonstrate results using three input images:


Another major update in Qwen-Image-Edit-2509 is enhanced consistency.

First, regarding person consistency, Qwen-Image-Edit-2509 shows significant improvement over Qwen-Image-Edit. Below are examples generating various portrait styles:

For instance, changing a person’s pose while maintaining excellent identity consistency:

Leveraging this improvement along with Qwen-Image’s unique text rendering capability, we find that Qwen-Image-Edit-2509 excels at creating meme images:

Of course, even with longer text, Qwen-Image-Edit-2509 can still render it while preserving the person’s identity:

Person consistency is also evident in old photo restoration. Below are two examples:

Naturally, besides real people, generating cartoon characters and cultural creations is also possible:

Second, Qwen-Image-Edit-2509 specifically enhances product consistency. We find that the model can naturally generate product posters from plain-background product images:

Or even simple logos:

Third, Qwen-Image-Edit-2509 specifically enhances text consistency and supports editing font type, font color, and font material:

Moreover, the ability for precise text editing has been significantly enhanced:

It is worth noting that text editing can often be seamlessly integrated with image editing—for example, in this poster editing case:


The final update in Qwen-Image-Edit-2509 is native support for commonly used ControlNet image conditions, such as keypoint control and sketches:

License Agreement

Qwen-Image is licensed under Apache 2.0.

Citation

We kindly encourage citation of our work if you find it useful.

@misc{wu2025qwenimagetechnicalreport,
      title={Qwen-Image Technical Report},
      author={Chenfei Wu and Jiahao Li and Jingren Zhou and Junyang Lin and Kaiyuan Gao and Kun Yan and Sheng-ming Yin and Shuai Bai and Xiao Xu and Yilei Chen and Yuxiang Chen and Zecheng Tang and Zekai Zhang and Zhengyi Wang and An Yang and Bowen Yu and Chen Cheng and Dayiheng Liu and Deqing Li and Hang Zhang and Hao Meng and Hu Wei and Jingyuan Ni and Kai Chen and Kuan Cao and Liang Peng and Lin Qu and Minggang Wu and Peng Wang and Shuting Yu and Tingkun Wen and Wensen Feng and Xiaoxiao Xu and Yi Wang and Yichang Zhang and Yongqiang Zhu and Yujia Wu and Yuxuan Cai and Zenan Liu},
      year={2025},
      eprint={2508.02324},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2508.02324},
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour — ComfyUI is installed on it. Open ComfyUI through the tunnel, then Templates → Qwen Image Edit 2509: it loads ComfyUI's own build of this model — the files above.

# on your rented machine: pip install diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image

pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", torch_dtype=torch.bfloat16).to("cuda")
start = load_image("/workspace/in.png")
image = pipe(prompt="the same scene at golden hour", image=start).images[0]
image.save("/workspace/out.png")
Renting a GPU — connect, tunnels, Python
# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }

# the files of the template “Qwen Image Edit 2509” (28.8 GB)
get Comfy-Org/Qwen-Image_ComfyUI split_files/vae/qwen_image_vae.safetensors vae
get Comfy-Org/Qwen-Image_ComfyUI split_files/text_encoders/qwen_2.5_vl_7b_fp8_scaled.safetensors text_encoders
get Comfy-Org/Qwen-Image-Edit_ComfyUI split_files/diffusion_models/qwen_image_edit_2509_fp8_e4m3fn.safetensors diffusion_models
get lightx2v/Qwen-Image-Lightning Qwen-Image-Edit-2509/Qwen-Image-Edit-2509-Lightning-4steps-V1.0-bf16.safetensors loras

start-comfyui
Renting a GPU — connect, tunnels, ComfyUI
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms