Model reference · open weights

HunyuanVideo-Foley

Audio tencent Music / audio 1 build Its own licence terms 514 dl/mo

HunyuanVideo-Foley is an open-weight audio or speech model from tencent. HunyuanVideo-Foley (BF16) weighs 10.3 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

What it is

Released bytencent
TypeAudio & music
TaskMusic / audio
Runs withhunyuanvideo-foley
Released2025-08-21
Popularity514 downloads / month
Weights10.3 GB (HunyuanVideo-Foley (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for HunyuanVideo-Foley (BF16)

Weights 10.3 GB (file size) · overhead about 1.6 GB.

CardOne streamCounted
memory
RTX 3060 12 GBdoes not fit11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What tencent says about HunyuanVideo-Foley


👥 Authors

Sizhe Shan1,2* • Qiulin Li1,3* • Yutao Cui1 • Miles Yang1 • Yuehai Wang2 • Qun Yang3 • Jin Zhou1† • Zhao Zhong1

🏢 1Tencent Hunyuan • 🎓 2Zhejiang University • ✈️ 3Nanjing University of Aeronautics and Astronautics

Read the full model card

*Equal contribution • †Project lead


🔥🔥🔥 News

  • [2025.9.29] 🚀 HunyuanVideo-Foley-XL Model Release - Release XL-sized model with offload inference support, significantly reducing VRAM requirements.
  • [2025.8.28] 🌟 HunyuanVideo-Foley Open Source Release - Inference code and model weights publicly available.

✨ Key Highlights

🎭 Multi-scenario Sync High-quality audio synchronized with complex video scenes

🧠 Multi-modal Balance Perfect harmony between visual and textual information

🎵 48kHz Hi-Fi Output Professional-grade audio generation with crystal clarity


📄 Abstract

🚀 Tencent Hunyuan open-sources HunyuanVideo-Foley an end-to-end video sound effect generation model!

A professional-grade AI tool specifically designed for video content creators, widely applicable to diverse scenarios including short video creation, film production, advertising creativity, and game development.

🎯 Core Highlights

🎬 Multi-scenario Audio-Visual Synchronization Supports generating high-quality audio that is synchronized and semantically aligned with complex video scenes, enhancing realism and immersive experience for film/TV and gaming applications.

⚖️ Multi-modal Semantic Balance Intelligently balances visual and textual information analysis, comprehensively orchestrates sound effect elements, avoids one-sided generation, and meets personalized dubbing requirements.

🎵 High-fidelity Audio Output Self-developed 48kHz audio VAE perfectly reconstructs sound effects, music, and vocals, achieving professional-grade audio generation quality.

🏆 SOTA Performance Achieved

HunyuanVideo-Foley comprehensively leads the field across multiple evaluation benchmarks, achieving new state-of-the-art levels in audio fidelity, visual-semantic alignment, temporal alignment, and distribution matching - surpassing all open-source solutions!

📊 Performance comparison across different evaluation metrics - HunyuanVideo-Foley leads in all categories


🔧 Technical Architecture

📊 Data Pipeline Design

🔄 Comprehensive data processing pipeline for high-quality text-video-audio datasets

The TV2A (Text-Video-to-Audio) task presents a complex multimodal generation challenge requiring large-scale, high-quality datasets. Our comprehensive data pipeline systematically identifies and excludes unsuitable content to produce robust and generalizable audio generation capabilities.

🏗️ Model Architecture

🧠 HunyuanVideo-Foley hybrid architecture with multimodal and unimodal transformer blocks

HunyuanVideo-Foley employs a sophisticated hybrid architecture:

  • 🔄 Multimodal Transformer Blocks: Process visual-audio streams simultaneously
  • 🎵 Unimodal Transformer Blocks: Focus on audio stream refinement
  • 👁️ Visual Encoding: Pre-trained encoder extracts visual features from video frames
  • 📝 Text Processing: Semantic features extracted via pre-trained text encoder
  • 🎧 Audio Encoding: Latent representations with Gaussian noise perturbation
  • ⏰ Temporal Alignment: Synchformer-based frame-level synchronization with gated modulation

📈 Performance Benchmarks

🎬 MovieGen-Audio-Bench Results

Objective and Subjective evaluation results demonstrating superior performance across all metrics

🏆 MethodPQ ↑PC ↓CE ↑CU ↑IB ↑DeSync ↓CLAP ↑MOS-Q ↑MOS-S ↑MOS-T ↑
FoleyGrafter6.272.723.345.680.171.290.143.36±0.783.54±0.883.46±0.95
V-AURA5.824.303.635.110.231.380.142.55±0.972.60±1.202.70±1.37
Frieren5.712.813.475.310.181.390.162.92±0.952.76±1.202.94±1.26
MMAudio6.172.843.595.620.270.800.353.58±0.843.63±1.003.47±1.03
ThinkSound6.043.733.815.590.180.910.203.20±0.973.01±1.043.02±1.08
HunyuanVideo-Foley (ours)6.592.743.886.130.350.740.334.14±0.684.12±0.774.15±0.75

🎯 Kling-Audio-Eval Results

Comprehensive objective evaluation showcasing state-of-the-art performance

🏆 MethodFD_PANNs ↓FD_PASST ↓KL ↓IS ↑PQ ↑PC ↓CE ↑CU ↑IB ↑DeSync ↓CLAP ↑
FoleyGrafter22.30322.632.477.086.052.913.285.440.221.230.22
V-AURA33.15474.563.245.805.693.983.134.830.250.860.13
Frieren16.86293.572.957.325.722.552.885.100.210.860.16
MMAudio9.01205.852.179.595.942.913.305.390.300.560.27
ThinkSound9.92228.682.396.865.783.233.125.110.220.670.22
HunyuanVideo-Foley (ours)6.07202.121.898.306.122.763.225.530.380.540.24

🎉 Outstanding Results! HunyuanVideo-Foley achieves the best scores across ALL evaluation metrics, demonstrating significant improvements in audio quality, synchronization, and semantic alignment.


🚀 Quick Start

📦 Installation

🔧 System Requirements

  • CUDA: 12.4 or 11.8 recommended
  • Python: 3.8+
  • OS:

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms