Model reference · open weights

stable-video-diffusion-img2vid-xt-1-1

Video vdo Image→video 1 build Its own licence terms 3k dl/mo

stable-video-diffusion-img2vid-xt-1-1 is an open-weight video model from vdo. stable-video-diffusion-img2vid-xt-1-1 (FP32) weighs 4.5 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Stable Video Diffusion 1.1 Image-to-Video is a 1.5B parameter latent diffusion model developed by Stability AI for generating short video clips from a single still image.
  • It produces 25 frames at a resolution of 1024x576 and is intended for research purposes only.
  • The model has limitations including the inability to control generation via text, render legible text, or achieve perfect photorealism, and it is distributed under an "other" licence.

Summary of the vdo/stable-video-diffusion-img2vid-xt-1-1 model card, 2026-10-01

What it is

Released byvdo
Released2024-02-05
Parameters1.5B
VRAM4.5 GB for the weights

What it runs on

Memory and cards for stable-video-diffusion-img2vid-xt-1-1 (FP32)

4.5 GBweights, file size
3.0 GBbiggest part
537 MBruntime overhead
CardWeightsMemory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. ComfyUI is installed on it. Open ComfyUI through the tunnel and load the workflow from the model's card on Hugging Face; choose this model's file in its loader.

# on your rented machine: pip install diffusers transformers accelerate ftfy
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import export_to_video

pipe = DiffusionPipeline.from_pretrained("vdo/stable-video-diffusion-img2vid-xt-1-1", torch_dtype=torch.bfloat16).to("cuda")
frames = pipe(prompt="a drone shot over a forest at sunrise").frames[0]
export_to_video(frames, "/workspace/out.mp4", fps=16)
Renting a GPU: connect, tunnels, Python
# on your rented machine (the ssh line is on its page in the console)
# get REPO FILE FOLDER: one file into /workspace/models/FOLDER, where ComfyUI loads it from
get() { hf download "$1" "$2" --local-dir /workspace/hf-files && mkdir -p "/workspace/models/$3" && mv "/workspace/hf-files/$2" "/workspace/models/$3/$4"; }

# the model (4.5 GB)
get vdo/stable-video-diffusion-img2vid-xt-1-1 svd_xt_1_1.safetensors checkpoints

start-comfyui
Renting a GPU: connect, tunnels, ComfyUI
# on your computer, in a second terminal: ComfyUI in your browser at http://localhost:8188
# HOST and PORT are your machine's, from its page in the console
ssh -L 8188:localhost:8188 dev@HOST -p PORT
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms