Model reference · open weights
Ovis-Omni-Embedding is an open-weight embedding model from ATH-MaaS. Ovis-Omni-Embedding-3B (BF16) weighs 11.1 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.
Summary of the ATH-MaaS/Ovis-Omni-Embedding-3B model card, 2026-10-05
What it is
| Released by | ATH-MaaS |
|---|---|
| Released | 2026-09-20 |
| VRAM | 11.1 GB for the weights |
What it runs on
| Card | Runs | Memory |
|---|---|---|
| RTX 3060 12 GB | does not fit | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
From the model card
Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling any-to-any retrieval with a single encoder.
The model is initialized from Qwen2.5-Omni-3B. Rather than attaching separate modality-specific embedding towers, it retains the native text tokenizer, vision encoder, audio encoder, and shared Thinker backbone. The speech-generation Talker and language-modeling head are removed, and the final-layer hidden state at the last non-padding token is used directly as the retrieval embedding.
Ovis-Omni-Embedding-3B is designed for multimodal search, retrieval-augmented generation, recommendation, visual-document retrieval, video and audio search, and agentic retrieval over tools, interfaces, and memory.
Figure 2 from the technical report. Ovis-Omni-Embedding-3B encodes text, visual, and audio inputs as an interleaved token sequence processed by the Qwen2.5-Omni Thinker with TMRoPE. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.
An input is formatted with a retrieval instruction through the native processor and chat template. Text, visual, and acoustic tokens are then processed as one interleaved sequence by the shared causal Transformer. Qwen2.5-Omni's time-aligned multimodal rotary position embedding preserves temporal alignment between audio and video.
Figure 4 from the technical report. The training corpus spans text, images, video, audio, visual documents, and interleaved inputs. Homogeneous-source sampling forms each micro-batch from one dataset and deduplicates pooled candidates to create informative in-batch negatives without positive collisions.
Training follows three stages:
Figure 5 from the technical report. Embedding Distillation transfers similarity distributions from complementary experts, while the inference-time low-rank module combines a shared PCA basis with lightweight residual adapters for compact embeddings.
Ovis-Omni-Embedding-3B is a bi-encoder, not a cross-encoder:
No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, audio clips, and multimodal items are all treated as candidates in the same embedding space.
MMEB-v3 is an omni-modal benchmark comprising 190 datasets across image, video, visual-document, text, audio, and agent retrieval. Ovis-Omni-Embedding-3B achieves 58.46 overall, outperforming the strongest compared baseline by 5.19 points, and ranks first on the aggregate score of every modality group.
| Group | Ovis-Omni-Embedding-3B | Best compared baseline | Margin |
|---|---|---|---|
| Image | 77.55 | 73.83 | +3.72 |
| Video | 64.99 | 59.37 | +5.62 |
| Visual document | 78.26 | 75.37 | +2.89 |
| Text | 47.15 | 43.62 | +3.53 |
| Audio | 50.08 | 43.17 | +6.91 |
| Agent | 45.52 | 39.42 | +6.10 |
| All 190 datasets | 58.46 | 53.27 | +5.19 |
Across the 31 aggregate and sub-task entries in the complete comparison, Ovis-Omni-Embedding-3B ranks first on 22 and second on 8. MultiConIR is the only entry on which it falls outside the top two.
Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 190 MMEB-v3 datasets. MMEB-v3 primarily uses Hit@1 for image, video, audio, and agent tasks and nDCG@5 for text and visual-document retrieval.
| Benchmark | Ovis-Embedding-Omni-3B | Best compared baseline | Evaluation scope |
|---|---|---|---|
| MAEB (beta) | 57.29 | LCO-Embedding-Omni-7B: 53.54 | Mean over 30 audio embedding tasks |
| MVEB (beta) | 61.77 | LCO-Embedding-Omni-7B: 57.58 | Mean over 23 video and audio-video embedding tasks |
| RTEB | 67.35 | Qwen3-Embedding-4B: 67.27 | 15-task English public retrieval split |
These benchmark families use their own official aggregation procedures, so their scores should not be averaged together. MAEB and MVEB results are local evaluations inserted into the corresponding leaderboard snapshots, as described in the technical report.
The native output width is 2048 because no embedding projection head is added to the backbone. For deployments with tighter storage or latency budgets, the post-hoc elastic-dimension module supports 1024, 512, 256, and 128 dimensions. It combines a shared, modality-balanced PCA rotation with a zero-initialized residual linear adapter and fo
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.