Model reference · open weights

Ovis-VL-Embedding

Embeddings ATH-MaaS Embeddings 1 build Open weights 513 dl/mo

Ovis-VL-Embedding is an open-weight embedding model from ATH-MaaS. Ovis-VL-Embedding-2B (BF16) weighs 4.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Ovis-VL-Embedding-2B is a compact vision-language embedding model developed by ATH-MaaS for feature extraction and cross-modal retrieval.
  • It processes text, images, visual documents, and video into a single 2048-dimensional representation space, supporting a context length of 262,144 tokens.
  • The model is released under the Apache 2.0 license and is designed for applications such as multimodal RAG, image search, and recommendation systems.

Summary of the ATH-MaaS/Ovis-VL-Embedding-2B model card, 2026-10-04

What it is

Released byATH-MaaS
Released2026-09-21
VRAM4.4 GB for the weights

What it runs on

Memory and cards for Ovis-VL-Embedding-2B (BF16)

4.4 GBweights, file size
1.1 GBruntime overhead
CardRunsMemory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

From the model card

What ATH-MaaS says about Ovis-VL-Embedding

Read the model card

Ovis-VL-Embedding-2B is a compact vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling cross-modal retrieval with a single 2B-scale encoder.

The model is initialized from Qwen3.5-2B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added.

Ovis-VL-Embedding-2B is designed for efficient multimodal search, multimodal RAG, image and visual-document retrieval, video search, recommendation, and nearest-neighbor matching under constrained serving budgets.

Architecture

Figure 2 from the technical report. Ovis-VL-Embedding-2B encodes text, images, visual documents, and sampled video frames as one interleaved sequence. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.

The Qwen3.5-2B backbone contains 24 language layers with hidden size 2048. Its hybrid stack repeats three Gated DeltaNet layers followed by one gated full-attention layer, combining efficient long-context processing with periodic global token interaction. Multimodal positional encoding preserves temporal and two-dimensional spatial coordinates for visual tokens.

Training

Training follows three stages:

  1. Multimodal contrastive pretraining. Large-scale, multi-task data establish broad alignment across text, image, document, and video inputs. The objective combines difficulty-aware focal contrastive learning with similarity-distribution distillation.
  2. Full-parameter homogeneous finetuning. High-quality downstream data refine fine-grained discrimination. Each micro-batch is drawn from one dataset so that gathered candidates form task-consistent negatives.
  3. Annealing Embedding Distillation. Teacher-correct examples are retained, unresolved student examples are emphasized, and confidence-adaptive forward-KL supervision transfers complementary expert capabilities.

Retrieval interface

Ovis-VL-Embedding-2B is a bi-encoder, not a cross-encoder:

  1. Pair the query with the task instruction and format it with the native processor and chat template.
  2. Encode queries and candidates independently.
  3. Extract the final-layer hidden state at the last non-padding token.
  4. L2-normalize the 2048-dimensional query and candidate embeddings.
  5. Rank candidates by cosine similarity, equivalently the dot product of the normalized vectors.

No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, and interleaved multimodal items are all treated as candidates in the same embedding space.

Performance

MMEB-v2

MMEB-v2 evaluates vision-language embeddings over 78 datasets spanning image, video, and visual-document tasks. Ovis-VL-Embedding-2B achieves 77.46 overall, outperforming the strongest compared baseline by 2.04 points.

GroupOvis-VL-Embedding-2BBest compared baselineResult
Image80.6277.41+3.21
Video67.1268.84-1.72
Visual document80.4779.86+0.61
All 78 datasets77.4675.42+2.04

The model ranks first on all four image sub-tasks, video classification, video moment retrieval, the visual-document aggregate, ViDoRe-V1, and out-of-distribution visual-document retrieval. Its strongest gains come from image understanding and document retrieval, while its compact scale retains competitive temporal-video performance.

Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 78 MMEB-v2 datasets.

Embedding dimensions

The native output width is 2048, inherited directly from the Qwen3.5-2B backbone because no embedding projection head is added. Queries and candidates must use the same preprocessing, pooling rule, dimensionality, and L2 normalization.

Intended use

The model is intended for embedding extraction and retrieval over supported unimodal or interleaved multimodal content, including:

  • semantic text and cross-modal search;
  • text-to-image, image-to-text, and image-to-image retrieval;
  • multimodal RAG indexing and retrieval;
  • visual-document and page retrieval;
  • text-to-video and video retrieval;
  • recommendation and nearest-neighbor matching.

Limitations

  • This checkpoint does not natively support audio input. Use Ovis-Embedding-Omni-3B for audio and general omni-modal retrieval.
  • This checkpoint produces retrieval embeddings; it is not intended as a text or image generation model.
  • Retrieval quality depends on task-appropriate query instructions and the native preprocessing and chat template.
  • Performance varies by task. Video question answering, general video retrieval, and ViDoRe-V2 remain below the strongest specialist baselines in the reported comparison.
  • Benchmark scores may not directly predict performance on a new domain. Evaluate with representative queries, candidates, and retrieval metrics before deployment.

Citation

If you find our embedding models useful, please consider citing our technical report:

@article{ovisembedding2026,
  title   = {Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings},
  author  = {{Ovis-Embedding Team}},
  journal = {arXiv preprint arXiv:2609.25165},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.25165}
}

License

This model is released under the Apache 2.0 license.

Resources

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms