Model reference · open weights

Cosmos3-Super-Image2Video

Cosmos3-Super-Image2Video is an open-weight video model from nvidia, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required Video nvidia 1 variants 80k downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What Cosmos3-Super-Image2Video is

Cosmos 3: Omnimodal World Models for Physical AI Model Collection | Code | White Paper | Website NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications. Model Overview: Cosmos3-Super-Image2Video Description Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning. This model is ready for commercial and non-commercial use. Model Developer: NVIDIA Model Versions - Cosmos3-Nano: - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications. - Cosmos3-Super: - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications. - Cosmos3-Nano-Policy-DROID: - Given language instructions and visual observations from the DROID robot platform, generate robot action trajectories for manipulation and control tasks. - Cosmos3-Super-Image2Video: - Given one input image and text instructions, generate temporally coherent video sequences that are consistent with the provided visual content. - Cosmos3-Super-Text2Image: - Given text input, generate high-fidelity images that are consistent with the provided description. License This model is released under the OpenMDW1.1 Deployment Geography Global Use Case Physical AI: Encompassing robotics, autonomous vehicles (AV), and smart space environments, inc

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makernvidia
TypeVideo models
Parameters (lead)64.6B
Variants1
Runs withcosmos
Released2026-05-21
Popularity80k downloads / month
Likes155
LicenceCommercial licence needed

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Cosmos3-Super-Image2Video64.6BBF16~148.6 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys cosmos3-super-image2video for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (cosmos3-super-image2video below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/videos/generations \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"cosmos3-super-image2video","prompt":"a drone shot over a forest"}'

Details

Languages, data & research

Tags

cosmos diffusers safetensors cosmos3_omni nvidia cosmos3 vllm-omni sglang sglang-diffusion image-to-video video-generation

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want Cosmos3-Super-Image2Video on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More video models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms