Model reference · open weights
TBDub is an open-weight video model from TaoLiveAIGC. TBDub (BF16) weighs 25.2 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.
What it is
| Released by | TaoLiveAIGC |
|---|---|
| Type | Video models |
| Task | Video edit |
| Released | 2026-09-04 |
| Popularity | 501 downloads / month |
| Weights | 25.2 GB (TBDub (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 25.2 GB (file size) · its biggest part 12.6 GB · overhead about 537 MB.
| Card | The weights | Counted memory |
|---|---|---|
| RTX 3060 12 GB | does not fit | 11.6 GB |
| RTX 4060 Ti 16 GB | tight (encoders offloaded) | 15.4 GB |
| RTX 3090 24 GB | fits (encoders offloaded) | 23.4 GB |
| RTX 4090 24 GB | fits (encoders offloaded) | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; a video's working memory grows with its resolution and length and is not estimated yet. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
High-quality lip sync. Fast, multilingual, and robust.
Give an existing video a new voice. TBDub synchronizes the speaker's lips with new speech while preserving their appearance, motion, and background.
Input: source video + driving audio → Output: lip-synced video.
Code · Teacher · Student — all open source under Apache-2.0.
One voice track, four faces — multilingual lip sync with the two-step Student Model. Watch examples in seven languages across microphone and on-screen text occlusions, large head turns, and rapid head motion. The video opens with the original clips, then shows the generated results. Turn on sound to hear the language changes.
Open the video · Run the Student Model
V1.1 reduces checkpoint download sizes and adds an optional lower-memory mode to the original V1.0 release. It uses the same trained Teacher and Student models; the paper's quality and H20 speed results remain the original evaluation.
--cpu-offload if inference runs out
of GPU memory. In this mode, inactive models stay in system RAM.
Measured process peaks were 19.49 GiB for Teacher and 15.01 GiB
for Student at 512×512 and 376 output frames on an RTX PRO 5000 72GB.
See the measured workload and tradeoffs below; these are separate from
the paper's H20 speed benchmark.V1.1 release notes · Project update and memory measurements
MOS: 38 TalkVid clips, 114 ratings per method, on a 0–5 scale. Speed: VAE encode to decode, excluding preprocessing, audio encoding, and file output. Full evaluation protocol.
Watch comparisons · Get the code · Download models
| File | Description |
|---|---|
tbdub_teacher.safetensors | Complete BF16 Teacher for standard 30-step inference |
tbdub_student.safetensors | Complete BF16 Student for two-step inference |
null_prompt_emb.pt | Null text embedding used during inference |
Choose one variant. Each checkpoint contains the complete DiT model and uses the same runtime parameter values as V1.0. Both variants also need the shared VAE, HuBERT, fixed prompt embedding, and Face Landmarker files listed below. For either variant, the complete required model download is about 15.27 GB: 12.59 GB for its DiT checkpoint plus about 2.68 GB of shared auxiliary files. The other variant's checkpoint is unnecessary. Sizes here use decimal GB and exclude Python packages and CUDA libraries.
For standard 30-step inference, download tbdub_teacher.safetensors. Its
12.59 GB DiT download replaces the 22.87 GB required by the V1.0 setup.
Exact size: 12,591,523,040 bytes. Teacher SHA-256:
e73483dfde3b960d3d22f85e356c378067ccd3cff81f3b668d92167b8de605d0.
For two-step inference, download tbdub_student.safetensors. Its BF16 file is
12,591,523,048 bytes (about 12.59 GB), compared with the 37.77 GB DiT download
documented for V1.0. All 1,245 tensors match the original Student cast to BF16,
the dtype already used by inference. The storage change reduces download
size; the optional CPU offload described below reduces GPU memory use.
SHA-256: 4e86900f9744b3650bfe58451ae89b3afba758233a3edf5947516268c1ea11e6.
The manifest lists each variant's single checkpoint, its checksum, and the shared auxiliary dependencies.
The inference code and complete installation instructions are available in the GitHub repository:
Clone the repository:
git clone https://github.com/TaoLiveAIGC/TBDub.git
cd TBDub
Use Linux with an NVIDIA GPU, Python 3.10 (validated environment), and ffmpeg on PATH. Install a
CUDA-compatible PyTorch/torchvision build in your environment, then install
the project dependencies:
pip install -r requirements.txt
Keep only the pinned opencv-contrib-python package;
do not install opencv-python alongside it because both provide cv2.
Download the configuration manifest together with the model variant you plan to use. The current GitHub inference entry point reads checkpoints/config.json, validates its runtime compatibility, and uses it
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
Running it yourself
Rent a machine by the hour — ComfyUI is installed on it. Open ComfyUI through the tunnel and load the workflow from the model's card on Hugging Face; choose this model's file in its loader.
# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }
# the model (11.7 GB)
get TaoLiveAIGC/TBDub tbdub_teacher.safetensors diffusion_models
start-comfyui
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT