Model reference · open weights
MoonViT-SO is an open-weight embedding model from moonshotai. MoonViT-SO-400M (BF16) weighs 834 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | moonshotai |
|---|---|
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 417M |
| Runs with | transformers |
| Released | 2025-04-10 |
| Popularity | 807 downloads / month |
| Weights | 834 MB (MoonViT-SO-400M (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 834 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
MoonViT is a Native-resolution Vision Encoder, which is initialized from and continually pre-trained on SigLIP-SO-400M. To facilitate the standalone use of MoonViT, we have separated the implementation and weights of MoonViT from moonshotai/Kimi-VL-A3B-Instruct.
If you are interested in the training process of MoonViT, you are welcome to read Paper Kimi-VL Technical Report.
from PIL import Image
from transformers import AutoModel, AutoImageProcessor
model_path = "moonshotai/MoonViT-SO-400M"
model = AutoModel.from_pretrained(
model_path,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
processor = AutoImageProcessor.from_pretrained(model_path, trust_remote_code=True)
image_path = "./figures/demo.png"
image = Image.open(image_path)
images_processed = processor(image, return_tensors="pt").to(dtype=model.dtype, device=model.device)
image_features: list = model(images_processed.pixel_values, images_processed.image_grid_hws)
print(f"dtype: {image_features[0].dtype}, shape: {image_features[0].shape}")
# dtype: torch.bfloat16, shape: torch.Size([1092, 4, 1152])
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.