Model reference · open weights

HunyuanImage-3-verbatim-flashpack

NEW · this week Image fal Image edit 1 build Its own licence terms 0 dl/mo

HunyuanImage-3-verbatim-flashpack is an open-weight image model from fal.

  • HunyuanImage-3-verbatim-flashpack is an image-to-image model developed by fal that uses a native multimodal architecture for creative editing and multi-image fusion.
  • It features a Mixture of Experts design with 80 billion total parameters and 13 billion activated per token, supporting a context length of 22,800 tokens.
  • The model is distributed under an "other" licence.

Summary of the fal/HunyuanImage-3-Instruct-verbatim-flashpack model card, 2026-10-02

What it is

Released byfal
Released2026-10-01

From the model card

What fal says about HunyuanImage-3-verbatim-flashpack

Read the model card

中文文档

🎨 HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation

💻 Official website(官网) Try our model!&nbsp&nbsp

🔥🔥🔥 News

🧩 Community Contributions

If you develop/use HunyuanImage-3.0 in your projects, welcome to let us know.

📑 Open-source Plan

  • HunyuanImage-3.0 (Image Generation Model)
    • [x] Inference
    • [x] HunyuanImage-3.0 Checkpoints
    • [x] HunyuanImage-3.0-Instruct Checkpoints (with reasoning)
    • [x] vLLM Support
    • [x] Distilled Checkpoints
    • [x] Image-to-Image Generation
    • [ ] Multi-turn Interaction

🗂️ Contents


📖 Introduction

HunyuanImage-3.0 is a groundbreaking native multimodal model that unifies multimodal understanding and generation within an autoregressive framework. Our text-to-image and image-to-image model achieves performance comparable to or surpassing leading closed-source models.

✨ Key Features

  • 🧠 Unified Multimodal Architecture: Moving beyond the prevalent DiT-based architectures, HunyuanImage-3.0 employs a unified autoregressive framework. This design enables a more direct and integrated modeling of text and image modalities, leading to surprisingly effective and contextually rich image generation.

  • 🏆 The Largest Image Generation MoE Model: This is the largest open-source image generation Mixture of Experts (MoE) model to date. It features 64 experts and a total of 80 billion parameters, with 13 billion activated per token, significantly enhancing its capacity and performance.

  • 🎨 Superior Image Generation Performance: Through rigorous dataset curation and advanced reinforcement learning post-training, we've achieved an optimal balance between semantic accuracy and visual excellence. The model demonstrates exceptional prompt adherence while delivering photorealistic imagery with stunning aesthetic quality and fine-grained details.

  • 💭 Intelligent Image Understanding and World-Knowledge Reasoning: The unified multimodal architecture endows HunyuanImage-3.0 with powerful reasoning capabilities. It under stands user's input image, and leverages its extensive world knowledge to intelligently interpret user intent, automatically elaborating on sparse prompts with contextually appropriate details to produce superior, more complete visual outputs.

🚀 Usage

📦 Environment Setup

  • 🐍 Python: 3.12+ (recommended and tested)
  • ⚡ CUDA: 12.8
📥 Install Dependencies
# 1. First install PyTorch (CUDA 12.8 Version)
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128

# 2. Install tencentcloud-sdk for Prompt Enhanceme

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. How to run this model is on its model card.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms