Model reference · open weights

TBDub

Video TaoLiveAIGC Video edit 1 build Open weights 501 dl/mo

TBDub is an open-weight video model from TaoLiveAIGC. TBDub (BF16) weighs 25.2 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

What it is

Released byTaoLiveAIGC
TypeVideo models
TaskVideo edit
Released2026-09-04
Popularity501 downloads / month
Weights25.2 GB (TBDub (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for TBDub (BF16)

Weights 25.2 GB (file size) · its biggest part 12.6 GB · overhead about 537 MB.

CardThe weightsCounted
memory
RTX 3060 12 GBdoes not fit11.6 GB
RTX 4060 Ti 16 GBtight (encoders offloaded)15.4 GB
RTX 3090 24 GBfits (encoders offloaded)23.4 GB
RTX 4090 24 GBfits (encoders offloaded)23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; a video's working memory grows with its resolution and length and is not estimated yet. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What TaoLiveAIGC says about TBDub

High-quality lip sync. Fast, multilingual, and robust.

Give an existing video a new voice. TBDub synchronizes the speaker's lips with new speech while preserving their appearance, motion, and background.

Input: source video + driving audio → Output: lip-synced video.

Code · Teacher · Student — all open source under Apache-2.0.

Read the full model card

See TBDub in action · 46-second demo

One voice track, four faces — multilingual lip sync with the two-step Student Model. Watch examples in seven languages across microphone and on-screen text occlusions, large head turns, and rapid head motion. The video opens with the original clips, then shows the generated results. Turn on sound to hear the language changes.

Open the video · Run the Student Model

V1.1 — Inference and deployment update

V1.1 reduces checkpoint download sizes and adds an optional lower-memory mode to the original V1.0 release. It uses the same trained Teacher and Student models; the paper's quality and H20 speed results remain the original evaluation.

  • Smaller downloads: each variant needs one complete BF16 DiT checkpoint of about 12.59 GB, plus the shared auxiliary files listed below. Download only the Teacher or Student you plan to use.
  • Optional lower GPU memory use: add --cpu-offload if inference runs out of GPU memory. In this mode, inactive models stay in system RAM. Measured process peaks were 19.49 GiB for Teacher and 15.01 GiB for Student at 512×512 and 376 output frames on an RTX PRO 5000 72GB. See the measured workload and tradeoffs below; these are separate from the paper's H20 speed benchmark.

V1.1 release notes · Project update and memory measurements

Why TBDub?

  • 🏆 Leading perceptual quality. Highest mean opinion scores in all three dimensions among the open-source methods in our study: 3.85 lip sync, 3.80 visual quality, and 3.78 identity preservation. Student leads the first two; Teacher leads identity. See results.
  • ⚡ Fast, two-step generation. 7.13 FPS on a single NVIDIA H20 at 512 × 512, or 13.93× the Teacher's core generation speed. See benchmark.
  • 🌍 Multilingual lip sync. Lip sync across languages, with examples including English, Chinese, Japanese, Korean, and Russian. See lip sync and identity preservation in multilingual reconstruction. Watch examples.
  • 🛡️ Robust in challenging scenes. Stable lip sync and appearance through large head turns, hands and microphones over the mouth, and rapid motion. Head turns · Occlusions.

MOS: 38 TalkVid clips, 114 ratings per method, on a 0–5 scale. Speed: VAE encode to decode, excluding preprocessing, audio encoding, and file output. Full evaluation protocol.

Watch comparisons · Get the code · Download models

Model Variants

FileDescription
tbdub_teacher.safetensorsComplete BF16 Teacher for standard 30-step inference
tbdub_student.safetensorsComplete BF16 Student for two-step inference
null_prompt_emb.ptNull text embedding used during inference

Choose one variant. Each checkpoint contains the complete DiT model and uses the same runtime parameter values as V1.0. Both variants also need the shared VAE, HuBERT, fixed prompt embedding, and Face Landmarker files listed below. For either variant, the complete required model download is about 15.27 GB: 12.59 GB for its DiT checkpoint plus about 2.68 GB of shared auxiliary files. The other variant's checkpoint is unnecessary. Sizes here use decimal GB and exclude Python packages and CUDA libraries.

For standard 30-step inference, download tbdub_teacher.safetensors. Its 12.59 GB DiT download replaces the 22.87 GB required by the V1.0 setup.

Exact size: 12,591,523,040 bytes. Teacher SHA-256: e73483dfde3b960d3d22f85e356c378067ccd3cff81f3b668d92167b8de605d0.

For two-step inference, download tbdub_student.safetensors. Its BF16 file is 12,591,523,048 bytes (about 12.59 GB), compared with the 37.77 GB DiT download documented for V1.0. All 1,245 tensors match the original Student cast to BF16, the dtype already used by inference. The storage change reduces download size; the optional CPU offload described below reduces GPU memory use.

SHA-256: 4e86900f9744b3650bfe58451ae89b3afba758233a3edf5947516268c1ea11e6.

The manifest lists each variant's single checkpoint, its checksum, and the shared auxiliary dependencies.

Installation and Inference

The inference code and complete installation instructions are available in the GitHub repository:

  • GitHub: https://github.com/TaoLiveAIGC/TBDub

Clone the repository:

git clone https://github.com/TaoLiveAIGC/TBDub.git
cd TBDub

Use Linux with an NVIDIA GPU, Python 3.10 (validated environment), and ffmpeg on PATH. Install a CUDA-compatible PyTorch/torchvision build in your environment, then install the project dependencies:

pip install -r requirements.txt

Keep only the pinned opencv-contrib-python package; do not install opencv-python alongside it because both provide cv2.

Download the configuration manifest together with the model variant you plan to use. The current GitHub inference entry point reads checkpoints/config.json, validates its runtime compatibility, and uses it

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour — ComfyUI is installed on it. Open ComfyUI through the tunnel and load the workflow from the model's card on Hugging Face; choose this model's file in its loader.

# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }

# the model (11.7 GB)
get TaoLiveAIGC/TBDub tbdub_teacher.safetensors diffusion_models

start-comfyui
Renting a GPU — connect, tunnels, ComfyUI
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms