Model reference · open weights
VAREdit is an open-weight image model from HiDream-ai. VAREdit (BF16) weighs 11.4 GB; the smallest configuration that runs it is RTX 4090 24 GB.
What it is
| Released by | HiDream-ai |
|---|---|
| Type | Image models |
| Task | Image edit |
| Based on | FoundationVision/Infinity |
| Released | 2025-08-14 |
| Popularity | 0 downloads / month |
| Weights | 11.4 GB (VAREdit (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 11.4 GB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.
| Card | One 1024×1024 image | Counted memory |
|---|---|---|
| RTX 3060 12 GB … RTX 4060 Ti 16 GB | does not fit | |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
VAREdit is an advanced image editing model built on the Infinity models, designed for high-quality instruction-based image editing.
Try our online demos: 🤗VAREdit-8B-1024 and 🤗VAREdit-8B-512.
| Model Variant | Resolutions | HuggingFace Model | Time (H800) | VRAM (GB) |
|---|---|---|---|---|
| VAREdit-8B-512 | 512×512 | VAREdit-8B-512 | ~0.7s | 50.41 |
| VAREdit-8B-1024 | 1024×1024 | VAREdit-8B-1024 | ~1.99s | 50.41 |
Before starting, ensure you have:
git clone https://github.com/HiDream-ai/VAREdit.git
cd VAREdit
pip install -r requirements.txt
Download the VAREdit model checkpoints:
# Download from HuggingFace
git lfs install
git clone https://huggingface.co/HiDream-ai/VAREdit
from infer import load_model, generate_image
model_components = load_model(
pretrain_root="HiDream-ai/VAREdit",
model_path="HiDream-ai/VAREdit/8B-1024.pth",
model_size="8B",
image_size=1024
)
# Generate edited image
edited_image = generate_image(
model_components,
src_img_path="assets/test.jpg",
instruction="Add glasses to this girl and change hair color to red",
cfg=3.0, # Classifier-free guidance scale
tau=0.1, # Temperature parameter
seed=42 # Optional random seed
)
| Parameter | Description | Default |
|---|---|---|
cfg | Classifier-free guidance scale | 3.0 |
tau | Temperature for sampling | 0.1 |
seed | Random seed for reproducibility | -1 (random) |
VAREdit/
├── infer.py # Main inference script
├── infinity/ # Core model implementations
│ ├── models/ # Model architectures
│ ├── dataset/ # Data processing utilities
│ └── utils/ # Helper functions
├── tools/ # Additional tools and scripts
│ └── run_infinity.py # Model execution utilities
├── assets/ # Demo images and resources
└── README.md # This file
| Method | Size | EMU-Edit Bal. | PIE-Bench Bal. | Time (A800) |
|---|---|---|---|---|
| InstructPix2Pix | 1.1B | 2.923 | 4.034 | 3.5s |
| UltraEdit | 7.7B | 4.541 | 5.580 | 2.6s |
| OmniGen | 3.8B | 4.674 | 3.492 | 16.5s |
| AnySD | 2.9B | 3.129 | 3.326 | 3.4s |
| EditAR | 0.8B | 3.305 | 4.707 | 45.5s |
| ACE++ | 16.9B | 2.076 | 2.574 | 5.7s |
| ICEdit | 17.0B | 4.785 | 4.933 | 8.4s |
| VAREdit (256px) | 2.2B | 5.565 | 6.684 | 0.5s |
| VAREdit (512px) | 2.2B | 5.662 | 6.996 | 0.7s |
| VAREdit (512px) | 8.4B | 7.792 | 8.105 | 1.2s |
| VAREdit (1024px) | 8.4B | 7.379 | 7.688 | 3.9s |
Note: The released 8B models are trained longer and on more data, so the performances are better than that in the paper.
This project is licensed under the MIT License - see the LICENSE file for details.
If you use VAREdit in your research, please cite:
@article{varedit2025,
title={Visual Autoregressive Modeling for Instruction-Guided Image Editing},
author={Mao, Qingyang and Cai, Qi and Li, Yehao and Pan, Yingwei and Cheng, Mingyue and Yao, Ting and Liu, Qi and Mei, Tao},
journal={arXiv preprint},
year={2025}
}
Note: This project is under active development. Features and code may change.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works
Running it yourself
Rent a machine by the hour — ComfyUI is installed on it. Open ComfyUI through the tunnel, load the workflow from the model's card on Hugging Face, and choose this file in its text encoder loader.
# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }
# the model (8.8 GB)
get HiDream-ai/VAREdit flan-t5-xl/model-00001-of-00002.safetensors text_encoders
start-comfyui
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT