Model reference · open weights

Ovis-Omni-Embedding

Embeddings ATH-MaaS Embeddings 1 build Open weights 512 dl/mo

Ovis-Omni-Embedding is an open-weight embedding model from ATH-MaaS. Ovis-Omni-Embedding-3B (BF16) weighs 11.1 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

  • Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model developed by ATH-MaaS for feature extraction across text, images, video, audio, and visual documents.
  • It maps these diverse inputs into a single representation space to enable any-to-any retrieval tasks such as search and recommendation.
  • The model uses a native output width of 2048 dimensions and is released under the Apache 2.0 license.

Summary of the ATH-MaaS/Ovis-Omni-Embedding-3B model card, 2026-10-05

What it is

Released byATH-MaaS
Released2026-09-20
VRAM11.1 GB for the weights

What it runs on

Memory and cards for Ovis-Omni-Embedding-3B (BF16)

11.1 GBweights, file size
1.1 GBruntime overhead
CardRunsMemory
RTX 3060 12 GBdoes not fit11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

From the model card

What ATH-MaaS says about Ovis-Omni-Embedding

Read the model card

Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling any-to-any retrieval with a single encoder.

The model is initialized from Qwen2.5-Omni-3B. Rather than attaching separate modality-specific embedding towers, it retains the native text tokenizer, vision encoder, audio encoder, and shared Thinker backbone. The speech-generation Talker and language-modeling head are removed, and the final-layer hidden state at the last non-padding token is used directly as the retrieval embedding.

Ovis-Omni-Embedding-3B is designed for multimodal search, retrieval-augmented generation, recommendation, visual-document retrieval, video and audio search, and agentic retrieval over tools, interfaces, and memory.

Architecture

Figure 2 from the technical report. Ovis-Omni-Embedding-3B encodes text, visual, and audio inputs as an interleaved token sequence processed by the Qwen2.5-Omni Thinker with TMRoPE. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.

An input is formatted with a retrieval instruction through the native processor and chat template. Text, visual, and acoustic tokens are then processed as one interleaved sequence by the shared causal Transformer. Qwen2.5-Omni's time-aligned multimodal rotary position embedding preserves temporal alignment between audio and video.

Training

Figure 4 from the technical report. The training corpus spans text, images, video, audio, visual documents, and interleaved inputs. Homogeneous-source sampling forms each micro-batch from one dataset and deduplicates pooled candidates to create informative in-batch negatives without positive collisions.

Training follows three stages:

  1. Omni-modal contrastive pretraining. Globally mixed candidates and cross-device in-batch negatives establish broad alignment. The objective combines difficulty-aware focal contrastive learning with similarity-distribution distillation from complementary modality experts.
  2. Full-parameter homogeneous finetuning. High-quality downstream data refine fine-grained discrimination. Samples within each micro-batch come from one dataset, and candidate deduplication prevents positive collisions and false in-batch negatives.
  3. Annealing Embedding Distillation. Teacher-correct examples are retained, unresolved student examples are upsampled, and confidence-adaptive forward-KL supervision transfers complementary expert capabilities without adding inference-time towers.

Figure 5 from the technical report. Embedding Distillation transfers similarity distributions from complementary experts, while the inference-time low-rank module combines a shared PCA basis with lightweight residual adapters for compact embeddings.

Retrieval interface

Ovis-Omni-Embedding-3B is a bi-encoder, not a cross-encoder:

  1. Pair the query with the task instruction and format it with the model's native processor and chat template.
  2. Encode queries and candidates independently.
  3. Extract the final-layer hidden state at the last non-padding token.
  4. L2-normalize the query and candidate embeddings.
  5. Rank candidates by cosine similarity, equivalently the dot product of the normalized vectors.

No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, audio clips, and multimodal items are all treated as candidates in the same embedding space.

Performance

MMEB-v3

MMEB-v3 is an omni-modal benchmark comprising 190 datasets across image, video, visual-document, text, audio, and agent retrieval. Ovis-Omni-Embedding-3B achieves 58.46 overall, outperforming the strongest compared baseline by 5.19 points, and ranks first on the aggregate score of every modality group.

GroupOvis-Omni-Embedding-3BBest compared baselineMargin
Image77.5573.83+3.72
Video64.9959.37+5.62
Visual document78.2675.37+2.89
Text47.1543.62+3.53
Audio50.0843.17+6.91
Agent45.5239.42+6.10
All 190 datasets58.4653.27+5.19

Across the 31 aggregate and sub-task entries in the complete comparison, Ovis-Omni-Embedding-3B ranks first on 22 and second on 8. MultiConIR is the only entry on which it falls outside the top two.

Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 190 MMEB-v3 datasets. MMEB-v3 primarily uses Hit@1 for image, video, audio, and agent tasks and nDCG@5 for text and visual-document retrieval.

Additional benchmark results

BenchmarkOvis-Embedding-Omni-3BBest compared baselineEvaluation scope
MAEB (beta)57.29LCO-Embedding-Omni-7B: 53.54Mean over 30 audio embedding tasks
MVEB (beta)61.77LCO-Embedding-Omni-7B: 57.58Mean over 23 video and audio-video embedding tasks
RTEB67.35Qwen3-Embedding-4B: 67.2715-task English public retrieval split

These benchmark families use their own official aggregation procedures, so their scores should not be averaged together. MAEB and MVEB results are local evaluations inserted into the corresponding leaderboard snapshots, as described in the technical report.

Embedding dimensions

The native output width is 2048 because no embedding projection head is added to the backbone. For deployments with tighter storage or latency budgets, the post-hoc elastic-dimension module supports 1024, 512, 256, and 128 dimensions. It combines a shared, modality-balanced PCA rotation with a zero-initialized residual linear adapter and fo

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms