Model reference · open weights

PixelUMM

NEW · this week LLMs nvidia Image→text 1 build Non-commercial 0 dl/mo

PixelUMM is an open-weight language model from NVIDIA.

  • PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint understanding and generation across text, images, and video directly in pixel space.
  • It features a decoder-only Transformer architecture with a Qwen3-8B language backbone and 15,199,672,064 parameters.
  • The model supports text-to-image generation, image-conditioned text generation, video-conditioned understanding, and video generation.
  • The source code is licensed under Apache License 2.0, while the model checkpoint is restricted to non-commercial research or evaluation purposes under the NVIDIA One-Way Noncommercial License.

Summary of the nvidia/PixelUMM model card, 2026-10-02

What it is

Released byNVIDIA
Released2026-10-01

From the model card

What NVIDIA says about PixelUMM

Read the model card

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint understanding and generation across text, images, and video directly in pixel space.

PixelUMM is released for non-commercial research or evaluation purposes only.

Overview

Images are represented as pixel patches instead of embeddings from a separate pretrained vision encoder, allowing understanding and generation to share a Transformer-based representation.

PixelUMM supports:

  • text-to-image generation;
  • image-conditioned text generation;
  • video-conditioned understanding; and
  • video generation.

Model Architecture

  • Architecture: Decoder-only Transformer with raw-pixel patch embeddings and an iterative pixel-generation head.
  • Language backbone: Qwen/Qwen3-8B, revision b968826d9c46dd6066d109eabc6255188de91218.
  • Image representation: RGB images divided into 16-by-16 pixel patches.
  • Generation: Iterative denoising for image and video generation.
  • Parameters: 15,199,672,064.

Inputs and Outputs

PixelUMM accepts text, RGB images, and video frames. It produces text, RGB images, and video frames.

Installation

PixelUMM requires Linux, an NVIDIA GPU, a CUDA build of PyTorch, and FlashAttention. Install PyTorch and FlashAttention for your CUDA environment, then install the remaining dependencies.

Intended Use

PixelUMM is intended for researchers and developers studying unified multimodal modeling, pixel-space representation learning, multimodal understanding, and image or video generation. Example uses include:

  • studying shared representations across text, images, and video;
  • evaluating encoder-free multimodal architectures;
  • generating images or video from text prompts; and
  • producing text conditioned on images or video.

PixelUMM has not been validated for production or high-stakes use.

Limitations and Safety

PixelUMM may produce inaccurate, offensive, or otherwise inappropriate content. It may not follow prompts reliably and may produce semantically or temporally inconsistent outputs. Performance may vary across languages, visual domains, video duration, resolution, aspect ratio, and hardware.

Users should implement appropriate safety measures, including content filtering, abuse monitoring, and access controls. Users are responsible for model inputs and outputs, for obtaining the rights and permissions required for input content, and for complying with applicable laws and regulations.

License

The PixelUMM source code is licensed under the Apache License 2.0. Third-party notices are provided in THIRD_PARTY_LICENSES.md.

The PixelUMM model checkpoint is a separate artifact licensed under the NVIDIA One-Way Noncommercial License. Its use is limited to non-commercial research or evaluation purposes.

PixelUMM uses the Qwen3-8B base model, which is licensed under Apache-2.0, and includes modified code derived from BAGEL. Users are responsible for complying with all applicable upstream licenses and terms.

References

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms