Model reference · open weights
Ovis-VL-Embedding is an open-weight embedding model from ATH-MaaS. Ovis-VL-Embedding-2B (BF16) weighs 4.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the ATH-MaaS/Ovis-VL-Embedding-2B model card, 2026-10-04
What it is
| Released by | ATH-MaaS |
|---|---|
| Released | 2026-09-21 |
| VRAM | 4.4 GB for the weights |
What it runs on
| Card | Runs | Memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
From the model card
Ovis-VL-Embedding-2B is a compact vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling cross-modal retrieval with a single 2B-scale encoder.
The model is initialized from Qwen3.5-2B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added.
Ovis-VL-Embedding-2B is designed for efficient multimodal search, multimodal RAG, image and visual-document retrieval, video search, recommendation, and nearest-neighbor matching under constrained serving budgets.
Figure 2 from the technical report. Ovis-VL-Embedding-2B encodes text, images, visual documents, and sampled video frames as one interleaved sequence. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.
The Qwen3.5-2B backbone contains 24 language layers with hidden size 2048. Its hybrid stack repeats three Gated DeltaNet layers followed by one gated full-attention layer, combining efficient long-context processing with periodic global token interaction. Multimodal positional encoding preserves temporal and two-dimensional spatial coordinates for visual tokens.
Training follows three stages:
Ovis-VL-Embedding-2B is a bi-encoder, not a cross-encoder:
No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, and interleaved multimodal items are all treated as candidates in the same embedding space.
MMEB-v2 evaluates vision-language embeddings over 78 datasets spanning image, video, and visual-document tasks. Ovis-VL-Embedding-2B achieves 77.46 overall, outperforming the strongest compared baseline by 2.04 points.
| Group | Ovis-VL-Embedding-2B | Best compared baseline | Result |
|---|---|---|---|
| Image | 80.62 | 77.41 | +3.21 |
| Video | 67.12 | 68.84 | -1.72 |
| Visual document | 80.47 | 79.86 | +0.61 |
| All 78 datasets | 77.46 | 75.42 | +2.04 |
The model ranks first on all four image sub-tasks, video classification, video moment retrieval, the visual-document aggregate, ViDoRe-V1, and out-of-distribution visual-document retrieval. Its strongest gains come from image understanding and document retrieval, while its compact scale retains competitive temporal-video performance.
Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 78 MMEB-v2 datasets.
The native output width is 2048, inherited directly from the Qwen3.5-2B backbone because no embedding projection head is added. Queries and candidates must use the same preprocessing, pooling rule, dimensionality, and L2 normalization.
The model is intended for embedding extraction and retrieval over supported unimodal or interleaved multimodal content, including:
If you find our embedding models useful, please consider citing our technical report:
@article{ovisembedding2026,
title = {Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings},
author = {{Ovis-Embedding Team}},
journal = {arXiv preprint arXiv:2609.25165},
year = {2026},
url = {https://arxiv.org/abs/2609.25165}
}
This model is released under the Apache 2.0 license.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.