Model reference · open weights

HunyuanDiT

Image Tencent-Hunyuan Text→image 1 build Its own licence terms 202k dl/mo

HunyuanDiT is an open-weight image model from Tencent-Hunyuan. HunyuanDiT-v1.1-Diffusers-Distilled (FP32) weighs 7.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.

HunyuanDiT is a 1.5B parameter text-to-image diffusion transformer developed by Tencent-Hunyuan. It supports both English and Chinese prompts and is designed for 25-step image generation. The model is available under an other licence.

Summary of the Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled model card, 2026-10-01

What it is

Released byTencent-Hunyuan
TypeImage models
TaskText→image
Parameters (lead)1.5B
Runs withdiffusers
Released2024-06-14
Popularity202k downloads / month
Weights7.2 GB (HunyuanDiT-v1.1-Diffusers-Distilled (FP32), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for HunyuanDiT-v1.1-Diffusers-Distilled (FP32)

Weights 7.2 GB (file size) · its biggest part 3.3 GB · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GBfits (encoders offloaded)11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more); "encoders offloaded" means only the biggest part is on the card at once — diffusers' model offload, or ComfyUI unloading the text encoder. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Tencent-Hunyuan says about HunyuanDiT

Read the model card

[Arxiv] [project page] [github]

This repo contains the distilled Hunyuan-DiT in 🤗 Diffusers format.

It supports 25-step text-to-image generation.

Dependency

Please install PyTorch first, following the instruction in https://pytorch.org

Install the latest version of transformers with pip:

pip install --upgrade transformers

Then install the latest github version of 🤗 Diffusers with pip:

pip install git+https://github.com/huggingface/diffusers.git

Example Usage with 🤗 Diffusers

import torch
from diffusers import HunyuanDiTPipeline

pipe = HunyuanDiTPipeline.from_pretrained("Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled", torch_dtype=torch.float16)
pipe.to("cuda")

# You may also use English prompt as HunyuanDiT supports both English and Chinese
# prompt = "An astronaut riding a horse"
prompt = "一个宇航员在骑马"
image = pipe(prompt).images[0]

📈 Comparisons

In order to comprehensively compare the generation capabilities of HunyuanDiT and other models, we constructed a 4-dimensional test set, including Text-Image Consistency, Excluding AI Artifacts, Subject Clarity, Aesthetic. More than 50 professional evaluators performs the evaluation.

🎥 Visualization

  • Chinese Elements

  • Long Text Input

🔥🔥🔥 Tencent Hunyuan Bot

Welcome to Tencent Hunyuan Bot, where you can explore our innovative products in multi-round conversation!

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. Runs as it is with diffusers — on the machine, in Python.

# on your rented machine: pip install diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

pipe = DiffusionPipeline.from_pretrained("Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled", torch_dtype=torch.bfloat16).to("cuda")
image = pipe(prompt="a red bicycle on a cobbled street").images[0]
image.save("/workspace/out.png")
Renting a GPU — connect, tunnels, Python
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms