Model reference · open weights
SANA-Video_2.0_5B_720p_4step is an open-weight video model from Efficient-Large-Model, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.
About
SANA-Video 2.0 5B 720p — 4-Step Preview Research preview: This distilled checkpoint is an early T2V-only preview. For the original 50-step T2V + TI2V release, use SANA-Video2.05B720p. Project page · Online demo · Documentation · GitHub · Paper SANA-Video 2.0 is an efficient diffusion transformer for high-resolution video generation. This preview provides a full-model EMA checkpoint distilled for four-stage text-to-video generation at 720p. It combines gated bidirectional linear-attention layers with periodic dense softmax-attention anchors and shared Attention Residual aggregation. Model details Checkpoint lineage and format The selected checkpoint is the global-step-1000 DMD EMA export initialized from the SANA-Video 2.0 SFT model after merging the ReFL step-500 adapter. DMD then updates the full transformer; this release is therefore a full model, not a LoRA adapter. The checkpoint contains a statedictema tensor mapping only. It does not contain optimizer, learning-rate scheduler, gradient-scaler, or training-loop state. The inference entry point unwraps statedictema, removes an optional model. key prefix, and casts the transformer to BF16. Files - checkpoints/SANAVideo2.05B720p4step.pth: distilled EMA transformer - config.yaml: clean 5B source-tower config with the preview frame/FPS defaults - demo/: verified seed-4 MP4 and poster generated from this checkpoint - LICENSE: Apache License 2.0 Verified 4-step example This 1280 × 736, 81-frame, 16 FPS sample was generated from the released checkpoint with seed 4 and the exact command shown below. Prompt: In a cozy, vintage room adorned with floral wallpaper, a cartoon rooster sits comfortably in a floral-patterned armchair, sipping from a bottle of beer. The rooster, with its vibrant red comb and wattle, displays a range of expressions—smiling, nodding, and opening its beak wide in a cheerful manner. The setting includes wooden furniture and another beer bottle on the table, adding to the relaxed atmosphere. The camera captures the rooster from a close-up angle, emphasizing its animated movements and lively demeanor. Four-stage sampling contract This model does not use a truncated DPM-Solver trajectory. At each f
Summarised from the published model card. Read the full card on the HuggingFace links below.
Specifications
| Maker | Efficient-Large-Model |
|---|---|
| Type | Video models |
| Variants | 1 |
| Runs with | sana |
| Based on | Efficient-Large-Model/SANA-Video_2.0_5B_720p |
| Released | 2026-08-27 |
| Popularity | 2k downloads / month |
| Likes | 1 |
| Licence | Open weights |
How it works
Variants
Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.
| Variant | Params | Precision | VRAM | Fits 16 GB | Weights |
|---|---|---|---|---|---|
| SANA-Video_2.0_5B_720p_4step | — | BF16 | — | — | Weights ↗ |
Using it via the API
Once AxForge deploys sana-video-2-0-5b-720p-4step for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sana-video-2-0-5b-720p-4step below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/videos/generations \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"sana-video-2-0-5b-720p-4step","prompt":"a drone shot over a forest"}'
Licence
Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗
Explore