Model reference · open weights

Cosmos-1.0-Diffusion-Text2World

Video nvidia Text→video 1 build Its own licence terms 932 dl/mo

Cosmos-1.0-Diffusion-Text2World is an open-weight video model from NVIDIA. Cosmos-1.0-Diffusion-7B-Text2World (BF16) weighs 24.6 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

  • Cosmos-1.0-Diffusion-Text2World is a 7-billion parameter diffusion transformer model developed by NVIDIA for text-to-video generation.
  • It produces 5-second, 121-frame video clips at a default resolution of 1280x704 pixels and 24 frames per second, with configurable aspect ratios and frame rates between 12 and 40 fps.
  • The model is designed for physical AI development and is released under the NVIDIA Open Model License, which permits commercial use and the creation of derivative models.

Summary of the nvidia/Cosmos-1.0-Diffusion-7B-Text2World model card, 2026-10-01

What it is

Released byNVIDIA
Released2025-01-07
VRAM24.6 GB for the weights

What it runs on

Memory and cards for Cosmos-1.0-Diffusion-7B-Text2World (BF16)

24.6 GBweights, file size
14.5 GBbiggest part
537 MBruntime overhead
CardWeightsMemory
RTX 3060 12 GBdoes not fit11.6 GB
RTX 4060 Ti 16 GBtight (encoders offloaded)15.4 GB
RTX 3090 24 GBfits (encoders offloaded)23.4 GB
RTX 4090 24 GBfits (encoders offloaded)23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. How to run this model is on its model card.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms