Model reference · open weights

SANA-Video_2.0_5B_720p

Available as managed deployment Video Efficient-Large-Model Text→video 1 variants 407 dl/mo

SANA-Video_2.0_5B_720p is an open-weight video model from Efficient-Large-Model. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerEfficient-Large-Model
TypeVideo models
TaskText→video
Runs withsana
Released2026-08-20
Popularity407 downloads / month
LicenceOpen weights

About

What SANA-Video_2.0_5B_720p is

SANA-Video 2.0 is an efficient diffusion transformer for high-resolution video generation. This release provides the 5B-class checkpoint jointly post-trained for text-to-video (T2V) and text-image-to-video (TI2V) generation at 720p for about eight seconds.

The model combines gated bidirectional linear-attention layers with periodic dense softmax-attention anchors and shared Attention Residual aggregation. It uses Gemma 2 2B IT for text conditioning and the LTX 2.3 VAE contract with 128 latent channels and (8, 32, 32) temporal/spatial compression.

Model details

PropertyValue
ArchitectureSanaVideo2_5B
Parameters4,466,980,960 trainable model parameters (4.47B)
Transformer32 layers, hidden size 2,560
Attention75% gated linear attention, 25% dense softmax anchors
Attention ResidualsShared, timestep-independent aggregation every 8 layers
TasksText-to-video and text-image-to-video
Output bucket736 × 1280, 193 frames, 24 FPS (about 8 seconds)
Text encodergoogle/gemma-2-2b-it
VAELTX 2.3, 128 latent channels, stride (8, 32, 32)
Recommended inferenceBF16, CFG 8, flow shift 12, 50 steps, motion score 20
LicenseApache 2.0

The checkpoint is an inference artifact containing only the merged model state_dict. It does not contain optimizer, scheduler, training-state, or standalone LoRA tensors. The EMA base weights and ReFL post-training adapter were merged before release. Stored tensors retain their merged source dtypes; the official inference entry point casts the transformer to BF16.

Files

  • checkpoints/SANA_Video_2.0_5B_720p.pth: merged transformer checkpoint
  • config.yaml: matching SANA training and inference configuration
  • LICENSE: Apache License 2.0

Verified release example

This sample was generated from the public checkpoint with seed 4. The encoded result is 1280 × 736, 193 frames, 24 FPS, and 8.04 seconds long.

Prompt: In a cozy, vintage room adorned with floral wallpaper, a cartoon rooster sits comfortably in a floral-patterned armchair, sipping from a bottle of beer. The rooster, with its vibrant red comb and wattle, displays a range of expressions—smiling, nodding, and opening its beak wide in a cheerful manner. The setting includes wooden furniture and another beer bottle on the table, adding to the relaxed atmosphere. The camera captures the rooster from a close-up angle, emphasizing its animated movements and lively demeanor.

Inference

Support for this checkpoint is included in the main branch of NVlabs/Sana:

git clone https://github.com/NVlabs/Sana.git
cd Sana
bash environment_setup.sh sana
conda activate sana

Place the Diffusers-format LTX 2.3 VAE at output/pretrained_models/LTX-2.3-Diffusers/, or update vae.vae_pretrained in config.yaml.

Text-to-video

The command below is the exact command used for the verified release example:

bash inference_video_scripts/inference_sana_video.sh \
  --np 1 \
  --config configs/sana_video2/SanaVideo2_5B_720p.yaml \
  --model_path hf://Efficient-Large-Model/SANA-Video_2.0_5B_720p/checkpoints/SANA_Video_2.0_5B_720p.pth \
  --txt_file=asset/samples/sana_video2_5b_720p_demo.txt \
  --cfg_scale 8 \
  --flow_shift 12 \
  --step 50 \
  --fps 24 \
  --motion_score 20 \
  --seed 4 \
  --work_dir output/sana_video2_t2v_720p_demo

Text-image-to-video

Each line in asset/samples/sample_i2v.txt contains a prompt and an input-image path separated by ``.

bash inference_video_scripts/inference_sana_video.sh \
  --np 1 \
  --config configs/sana_video2/SanaVideo2_5B_720p.yaml \
  --model_path hf://Efficient-Large-Model/SANA-Video_2.0_5B_720p/checkpoints/SANA_Video_2.0_5B_720p.pth \
  --txt_file=asset/samples/sample_i2v.txt \
  --task=ltx \
  --cfg_scale 8 \
  --flow_shift 12 \
  --step 50 \
  --fps 24 \
  --motion_score 20 \
  --work_dir output/sana_video2_ti2v_720p

The default 720p bucket is 736 × 1280 because both spatial dimensions must be divisible by 32. Frame counts must satisfy (num_frames - 1) % 8 == 0.

Intended use

This model is intended for research, evaluation, and creative generation of short videos from text, with optional first-frame image conditioning. It can also serve as a starting point for domain-specific fine-tuning under the license terms.

The model is not intended to produce factual evidence, identify people, make high-impact automated decisions, or generate content that violates privacy, copyright, applicable law, or platform policies.

Limitations and bias

  • Generated motion, anatomy, text rendering, object permanence, and physical interactions may be inconsistent, especially for crowded or highly dynamic scenes.
  • Prompt following can degrade for long, ambiguous, or compositionally complex instructions.
  • Image-conditioned generation can drift from fine details in the source image.
  • Outputs can reflect social and cultural biases present in training data and in the separately loaded text encoder.
  • The model does not independently verify whether generated content is factual, safe, or free of third-party rights.

Users should review outputs before publication, disclose synthetic media where appropriate, and add safeguards suited to their application.

Resources

Citation

If you use SANA-Video, please cite the SANA-Video work linked from the project page. SANA-Video 2.0-specific citation informa

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys sana-video-2-0-5b-720p for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sana-video-2-0-5b-720p below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/videos/generations \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"sana-video-2-0-5b-720p","prompt":"a drone shot over a forest"}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms