Model reference · open weights

Omni-Embed-Mini

Available as managed deployment Licence fee Embeddings MBZUAI Embeddings 1 variants 0 dl/mo

Omni-Embed-Mini is an open-weight embedding model from MBZUAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byMBZUAI
TypeEmbedding models
TaskEmbeddings
Runs withtransformers
Released2026-09-04
Popularity0 downloads / month
LicenceCommercial licence needed

About

What Omni-Embed-Mini is

2.3B parameters. Text, speech, general audio, image, video, and visually-rich documents in a single shared cosine space, from a backbone that is never updated.

Project page | Code | Sibling model: Omni-Embed-Mini-0.9B

Read the full model card

What this is

Omni-Embed-Mini recasts cross-modal alignment as self-distillation through one shared frozen causal backbone. Every media sample is paired with a dense cascaded caption; the teacher target is the EOS-pooled embedding of that caption produced by the identical frozen backbone that processes the student input. Teacher and student therefore inhabit byte-identical geometry, and because no text-side parameter is ever updated, adding audio cannot degrade the inherited text and vision representations.

Only the audio projectors and small phased LoRA adapters on the audio encoders are trained. A Matryoshka SigLIP contrastive objective and an online hybrid hard-negative miner supply the contrastive signal. This variant keeps the Qwen3-VL native visual tower, so image, video, and document-page quality is inherited intact from the backbone while speech and audio are added on top.

Results

Six modalities, evaluated with the pipeline in the code repository.

ModalityBenchmarkMetric2.3B v12.3B v20.9B v10.9B v2
TextMTEB-v2 BEIR-8nDCG@1047.9449.57
SpeechMAEB (12 tasks)mean48.8643.28
AudioMAEB (10 tasks)mean33.4433.43
ImageMMEB-V2 (10 tasks)hit@164.8026.29
VideoMMEB-V2 (6 tasks)hit@155.1818.48
Vis-DocViDoRe-v3 (7 tasks)nDCG@558.1046.92

v1 and v2 are tagged revisions of this same repository, so revision="v1.0" pins the numbers in the v1 column.

Quick start

pip install "transformers>=5.0" torch pillow numpy soundfile huggingface_hub
import torch
from transformers import AutoModel, AutoProcessor

REPO = "MBZUAI/Omni-Embed-Mini-2.3B"

model = AutoModel.from_pretrained(
    REPO, revision="v1.0", trust_remote_code=True, dtype=torch.bfloat16,
).cuda().eval()
processor = AutoProcessor.from_pretrained(REPO, revision="v1.0", trust_remote_code=True)

# OpenAI-style multimodal messages: text, audio, image, video, doc, or any composition.
messages = [{"role": "user", "content": [
    {"type": "audio", "audio": "path/to/clip.wav"},
    {"type": "text",  "text":  "rain on a tin roof at night"},
]}]

inputs = processor.apply_chat_template(
    messages, role="passage", tokenize=True, return_tensors="pt",
).to("cuda")
with torch.no_grad():
    doc = model(**inputs).pooler_output          # (1, 2048), already L2-normalised

# Query side. role="query" applies the retrieval instruction template.
# text_recipe="chat" puts a text query in the same subspace as media documents;
# omit it only for pure text-to-text retrieval, which uses the native recipe.
query = processor.apply_chat_template(
    [{"role": "user", "content": "rain on a tin roof at night"}],
    role="query", tokenize=True, return_tensors="pt", text_recipe="chat",
).to("cuda")
with torch.no_grad():
    q = model(**query).pooler_output

print("cosine:", float(q @ doc.T))               # normalised, so dot == cosine

Matryoshka (truncatable) embeddings

model(**inputs, truncate_dim=N).pooler_output returns a prefix-truncated, renormalised vector. Slicing pooler_output yourself is not equivalent unless you renormalise after. Supported N: 128, 256, 512, 1024, 2048. Use a smaller N to cut index size at a modest recall cost.

Architecture

ComponentThis model
Backbone (frozen)Qwen/Qwen3-VL-Embedding-2B
Embedding dim2048
Visionnative Qwen3-VL visual tower (inside backbone/; no separate file)
Speech encoderopenai/whisper-small
Audio encodermispeech/dasheng-base
Matryoshka dims128, 256, 512, 1024, 2048
Videonative video path (video_as_images: false)

Trained parameters: audio projectors plus phase-2 LoRA adapters on the audio encoders (48 Whisper and 24 Dasheng adapted tensors, merged into the released weights). The backbone is frozen at every stage and ships unmodified apart from a vocabulary resize that adds the media placeholder tokens. There is no vision_encoder.pt, which is expected for this variant, whose vision weights live inside backbone/.

Files

config.json                 modeling_omni_embed.py      processing_omni_embed.py
configuration_omni_embed.py chat_template.jinja         processor_config.json
tokenizer.json              tokenizer_config.json
backbone/                   # frozen Qwen3-VL backbone (vocab-resized), safetensors
whisper_encoder.pt          dasheng_encoder.pt
projector_weights.pt        # projectors; LoRA already merged into the encoders

The .pt files are PyTorch pickles, and the custom modeling code requires trust_remote_code=True. Load only from a source you trust.

Citation

@inproceedings{omniembedmini2026,
  title     = {Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation},
  author    = {TBD},
  booktitle = {TBD},
  year      = {2026},
  note      = {Camera-ready in preparation}
}

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys omni-embed-mini for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (omni-embed-mini below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"omni-embed-mini","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms