Model reference · open weights
PixelUMM is an open-weight language model from NVIDIA.
Summary of the nvidia/PixelUMM model card, 2026-10-02
What it is
| Released by | NVIDIA |
|---|---|
| Released | 2026-10-01 |
From the model card
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint understanding and generation across text, images, and video directly in pixel space.
PixelUMM is released for non-commercial research or evaluation purposes only.
Images are represented as pixel patches instead of embeddings from a separate pretrained vision encoder, allowing understanding and generation to share a Transformer-based representation.
PixelUMM supports:
b968826d9c46dd6066d109eabc6255188de91218.PixelUMM accepts text, RGB images, and video frames. It produces text, RGB images, and video frames.
PixelUMM requires Linux, an NVIDIA GPU, a CUDA build of PyTorch, and FlashAttention. Install PyTorch and FlashAttention for your CUDA environment, then install the remaining dependencies.
PixelUMM is intended for researchers and developers studying unified multimodal modeling, pixel-space representation learning, multimodal understanding, and image or video generation. Example uses include:
PixelUMM has not been validated for production or high-stakes use.
PixelUMM may produce inaccurate, offensive, or otherwise inappropriate content. It may not follow prompts reliably and may produce semantically or temporally inconsistent outputs. Performance may vary across languages, visual domains, video duration, resolution, aspect ratio, and hardware.
Users should implement appropriate safety measures, including content filtering, abuse monitoring, and access controls. Users are responsible for model inputs and outputs, for obtaining the rights and permissions required for input content, and for complying with applicable laws and regulations.
The PixelUMM source code is licensed under the Apache License 2.0. Third-party notices are provided in THIRD_PARTY_LICENSES.md.
The PixelUMM model checkpoint is a separate artifact licensed under the NVIDIA One-Way Noncommercial License. Its use is limited to non-commercial research or evaluation purposes.
PixelUMM uses the Qwen3-8B base model, which is licensed under Apache-2.0, and includes modified code derived from BAGEL. Users are responsible for complying with all applicable upstream licenses and terms.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.
How it works