Model reference · open weights
C-RADIOv4-1D-H is an open-weight embedding model from NVIDIA. C-RADIOv4-1D-H (FP32) weighs 2.9 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | NVIDIA |
|---|---|
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 1.5B |
| Runs with | transformers |
| Based on | nvidia/C-RADIOv4-H |
| Released | 2026-05-29 |
| Popularity | 1k downloads / month |
| Weights | 2.9 GB (C-RADIOv4-1D-H (FP32), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 2.9 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This model performs visual feature extraction. Unlike standard Vision Transformers that produce a fixed 2D grid of patch features, RADIO1D compresses an input image into a compact, variable-length 1D sequence of tokens. The number of output tokens (from 1 up to 256) can be selected by the user at inference time, providing a continuous accuracy/efficiency trade-off. For example, an image can be summarized into a single token for retrieval, or expanded to 256 tokens for fine-grained tasks such as OCR.
RADIO1D was produced by fine-tuning C-RADIOv4-H using multi-teacher agglomerative distillation from:
The encoder integrates a learnable Patch Merging block (4× sequence-length reduction, 2× channel expansion) part-way through the network for efficiency, and a lightweight Vision Transformer decoder is used only during training to project the 1D tokens back into a 2D-compatible grid for teacher alignment. At inference, only the encoder runs.
This model is ready for commercial or non-commercial use.
GOVERNING TERMS: Use of this model is governed by the NVIDIA Open Model License Agreement.
Global
The embeddings generated by this model are expected to be used by a downstream application. The variable-length 1D token output makes RADIO1D especially well suited to:
Hugging Face: 07/01/2026 via RADIO Collection of Models.
Architecture Type: Neural Network Network Architecture: Vision Transformer with encoder-decoder for elastic 1D token generation Number of model parameters: ~1.14B (encoder, used at inference); ~314M additional decoder parameters used only during training
The RADIO1D-H encoder is built from a ViT-H/16 backbone. Image patches (16×16 pixels) are flattened into a 1D sequence and processed by 24 transformer blocks at embedding dimension 1280, followed by a learnable Patch Merging block that groups 2×2 neighboring tokens (reducing sequence length by 4× and expanding the channel dimension by ρ=2 to 2560), followed by 8 further transformer blocks at the wider dimension. During training, a length ℓ is sampled stochastically from a triangular distribution p(x)=2−2x, and only the first ℓ encoder tokens are retained (a form of nested dropout). Earlier tokens are therefore encouraged to encode global, high-level semantics while later tokens specialize in finer details.
A small ViT decoder (~314M parameters, 6 blocks with one Patch Splitting upscale) is used only during training. It receives duplicated learnable query tokens equal in number to the original patch grid plus the encoder's ℓ tokens as additional register tokens, and cross-attention reconstructs a 2D-compatible feature grid that is aligned to the teacher representations.
At inference, only the encoder is used and the user specifies the desired number of output tokens.
Input Type(s): Image Input Format(s): Red, Green, Blue (RGB) Input Parameters: Two Dimensional (2D) Other Properties Related to Input: Image resolutions up to 2048×2048 in increments of 16 pixels. Training used a mix of low-resolution images (128, 192, 224, 256, 384, 432 px) and high-resolution images (512, 768, 1024, 1152 px).
Output Type(s): Embeddings
Output Format: Tensor
Output Parameters: One Dimensional (1D) — variable-length sequence of tokens
Other Properties Related to Output: The encoder returns a sequence of prefix tokens (CLS + register tokens) followed by ℓ global 1D tokens, where ℓ is selected by the caller at inference (typical values: 1, 8, 32, 64, 128, 192, 224, 256). Tokens are ordered hierarchically: token 0 encodes the strongest global summary (e.g., 85.0% k-NN Top-1 on ImageNet-1k with a single token), and later tokens add progressively finer detail. A downstream model is required to leverage the image features. Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g., GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Runtime Engine(s):
Supported Hardware Microarchitecture Compatibility:
[Preferred/Supported] Operating System(s):
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-ca
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.