Model reference · open weights

JoyAI-Video-Edit

Video jdopensource Video edit 1 build Open weights 3k dl/mo

JoyAI-Video-Edit is an open-weight video model from jdopensource. JoyAI-Video-Edit (BF16) weighs 1.5 GB; the smallest configuration that runs it is RTX 3060 12 GB.

JoyAI-Video-Edit is a real-time, instruction-guided video editing system developed by jdopensource that processes live or uploaded video streams using natural-language instructions. The model features a 16B-parameter multimodal diffusion transformer and supports English and Chinese, operating under the Apache 2.0 license. It achieves 30 FPS end-to-end throughput at 720 × 1248 resolution and can run on a single GeForce RTX 5090 GPU at 840 × 480 resolution.

Summary of the jdopensource/JoyAI-Video-Edit model card, 2026-10-01

What it is

Released byjdopensource
TypeVideo models
TaskVideo edit
Runs withjoyai-video-edit
Released2026-08-04
Popularity3k downloads / month
Weights1.5 GB (JoyAI-Video-Edit (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for JoyAI-Video-Edit (BF16)

Weights 1.5 GB (file size) · overhead about 537 MB.

CardThe weightsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; a video's working memory grows with its resolution and length and is not estimated yet. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What jdopensource says about JoyAI-Video-Edit

Read the model card

Real-Time Open-Ended Video Editing with Autoregressive Diffusion

src="https://raw.githubusercontent.com/jd-opensource/JoyAI-Video-Edit/main/assets/teaser.jpg"
width="96%"
alt="JoyAI-Video-Edit teaser"

🐶 JoyAI-Video-Edit

JoyAI-Video-Edit is a real-time, instruction-guided video editing system for open-ended video streams. Given a live camera stream or uploaded video and a natural-language edit instruction, it edits frames causally as they arrive, without waiting for the full video, requiring a predefined video length, or revisiting future frames. In our deployment benchmark, the full end-to-end pipeline reaches 30 FPS at 720 × 1248, pushing video editing from offline batch processing toward interactive streaming generation.

The system combines an MLLM-based condition encoder, a causal video VAE, and a 16B-parameter multimodal diffusion transformer. It is trained and deployed as an autoregressive diffusion editor, then accelerated with aligned autoregressive distribution matching distillation, long-horizon optimization, bounded KV-state inference, and deployment-oriented scheduling to sustain high-throughput 720p editing while reducing train-inference mismatch and accumulated temporal drift.


🔥 News

  • 2026.08.24 — 🎉 Consumer GPU support landed: real-time streaming video editing on a single GeForce RTX 5090 (32 GB) at 840 × 480 @ 24 FPS. Deployment Guide

  • 2026.08.15 — 🎉 Live demo released — real-time streaming video editing on a single RTX PRO 6000 (Blackwell) GPU: 840 × 480 @ 24 FPS or 720p @ 16 FPS. Try HuggingFace Demo

  • 2026.08.14 — 🎉 Released an upgraded checkpoint with significantly stronger reference-image-guided video editing (RV2V), delivering better subject and identity preservation, more faithful reference conditioning, and improved temporal consistency across long streams. Grab the new DiT weights.

  • 2026.08.05 — 🎉 We released the model checkpoints, deployment code, and technical report.


💎 Highlights

  • Real-time open-ended editing. Edits live or uploaded videos as frames arrive, without requiring the full sequence upfront.
  • Diverse instruction control. Supports subject edits, local edits, background changes, style transfer, motion changes, and reference-guided editing.
  • Autoregressive diffusion design. Combines an MLLM condition encoder, causal video VAE, and MMDiT backbone for streaming video editing.
  • High-throughput 720p deployment. Reaches 30 FPS end-to-end throughput at 720 × 1248 with bounded KV-state inference and stable per-chunk compute.

📦 Model Files

FilePurpose
dit/joyai_video_edit_dit_0811.pthCurrent DiT weights — use this one (stronger RV2V, released 2026.08.14)
dit/joyai_video_edit_dit_0804.pthInitial release (kept for reproducibility; the server does not use it)
vae/Causal video VAE (config + safetensors)

Place them under deploy/deps/checkpoints/JoyAI-Video-Edit/ — see the Deployment Guide for the full layout and the exact hf download command.


🎬 Demo

Try our online real-time video editing demo:

👉 https://huggingface.co/spaces/wxDai/joyai-video-edit

Project repository:

👉 https://github.com/jd-opensource/JoyAI-Video-Edit

Technical report:

👉 https://arxiv.org/abs/2608.03974


📚 Citation

If JoyAI-Video-Edit is useful for your research or project, please cite:

@article{xiao2026joyai,
  title={JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion},
  author={Xiao, Yicheng and Dai, Wenxun and Qin, Xinran and Song, Lin and Zhang, Maoquan and Xu, Hang and Chen, Yukang and Li, Yitong and Zhang, Guohui and Zhang, Yuan and Zhang, Xuying and Zhang, Tommy and Yuan, Jianlong and Li, Peihao and Lu, Shuai and Fu, Siming and Zhao, Chuyang and Han, Xin and Huang, Jie and Li, Wenbo and Ma, Guoqing and Huang, Wei and Qi, Xiaojuan and Huang, Haoyang and Duan, Nan},
  journal={arXiv preprint arXiv:2608.03974},
  year={2026}
}

📄 License

JoyAI-Video-Edit is released under the Apache License 2.0.

Please refer to the project repository for the complete license:

https://github.com/jd-opensource/JoyAI-Video-Edit/blob/main/LICENSE

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour — ComfyUI is installed on it. Open ComfyUI through the tunnel, load the workflow from the model's card on Hugging Face, and choose this file in its Load VAE node.

# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }

# the model (1.4 GB)
get jdopensource/JoyAI-Video-Edit vae/diffusion_pytorch_model.safetensors vae

start-comfyui
Renting a GPU — connect, tunnels, ComfyUI
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms