Model reference · open weights
Omni-Embed-Mini is an open-weight embedding model from MBZUAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | MBZUAI |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Runs with | transformers |
| Released | 2026-09-04 |
| Popularity | 0 downloads / month |
| Licence | Commercial licence needed |
About
2.3B parameters. Text, speech, general audio, image, video, and visually-rich documents in a single shared cosine space, from a backbone that is never updated.
Omni-Embed-Mini recasts cross-modal alignment as self-distillation through one shared frozen causal backbone. Every media sample is paired with a dense cascaded caption; the teacher target is the EOS-pooled embedding of that caption produced by the identical frozen backbone that processes the student input. Teacher and student therefore inhabit byte-identical geometry, and because no text-side parameter is ever updated, adding audio cannot degrade the inherited text and vision representations.
Only the audio projectors and small phased LoRA adapters on the audio encoders are trained. A Matryoshka SigLIP contrastive objective and an online hybrid hard-negative miner supply the contrastive signal. This variant keeps the Qwen3-VL native visual tower, so image, video, and document-page quality is inherited intact from the backbone while speech and audio are added on top.
Six modalities, evaluated with the pipeline in the code repository.
| Modality | Benchmark | Metric | 2.3B v1 | 2.3B v2 | 0.9B v1 | 0.9B v2 |
|---|---|---|---|---|---|---|
| Text | MTEB-v2 BEIR-8 | nDCG@10 | 47.94 | 49.57 | ||
| Speech | MAEB (12 tasks) | mean | 48.86 | 43.28 | ||
| Audio | MAEB (10 tasks) | mean | 33.44 | 33.43 | ||
| Image | MMEB-V2 (10 tasks) | hit@1 | 64.80 | 26.29 | ||
| Video | MMEB-V2 (6 tasks) | hit@1 | 55.18 | 18.48 | ||
| Vis-Doc | ViDoRe-v3 (7 tasks) | nDCG@5 | 58.10 | 46.92 |
v1 and v2 are tagged revisions of this same repository, so revision="v1.0" pins the
numbers in the v1 column.
pip install "transformers>=5.0" torch pillow numpy soundfile huggingface_hub
import torch
from transformers import AutoModel, AutoProcessor
REPO = "MBZUAI/Omni-Embed-Mini-2.3B"
model = AutoModel.from_pretrained(
REPO, revision="v1.0", trust_remote_code=True, dtype=torch.bfloat16,
).cuda().eval()
processor = AutoProcessor.from_pretrained(REPO, revision="v1.0", trust_remote_code=True)
# OpenAI-style multimodal messages: text, audio, image, video, doc, or any composition.
messages = [{"role": "user", "content": [
{"type": "audio", "audio": "path/to/clip.wav"},
{"type": "text", "text": "rain on a tin roof at night"},
]}]
inputs = processor.apply_chat_template(
messages, role="passage", tokenize=True, return_tensors="pt",
).to("cuda")
with torch.no_grad():
doc = model(**inputs).pooler_output # (1, 2048), already L2-normalised
# Query side. role="query" applies the retrieval instruction template.
# text_recipe="chat" puts a text query in the same subspace as media documents;
# omit it only for pure text-to-text retrieval, which uses the native recipe.
query = processor.apply_chat_template(
[{"role": "user", "content": "rain on a tin roof at night"}],
role="query", tokenize=True, return_tensors="pt", text_recipe="chat",
).to("cuda")
with torch.no_grad():
q = model(**query).pooler_output
print("cosine:", float(q @ doc.T)) # normalised, so dot == cosine
model(**inputs, truncate_dim=N).pooler_output returns a prefix-truncated, renormalised
vector. Slicing pooler_output yourself is not equivalent unless you renormalise after.
Supported N: 128, 256, 512, 1024, 2048. Use a smaller N to cut index size at a modest recall cost.
| Component | This model |
|---|---|
| Backbone (frozen) | Qwen/Qwen3-VL-Embedding-2B |
| Embedding dim | 2048 |
| Vision | native Qwen3-VL visual tower (inside backbone/; no separate file) |
| Speech encoder | openai/whisper-small |
| Audio encoder | mispeech/dasheng-base |
| Matryoshka dims | 128, 256, 512, 1024, 2048 |
| Video | native video path (video_as_images: false) |
Trained parameters: audio projectors plus phase-2 LoRA adapters on the audio encoders
(48 Whisper and 24 Dasheng adapted tensors, merged into the released weights). The backbone is
frozen at every stage and ships unmodified apart from a vocabulary resize that adds the media
placeholder tokens. There is no vision_encoder.pt, which is expected for this variant,
whose vision weights live inside backbone/.
config.json modeling_omni_embed.py processing_omni_embed.py
configuration_omni_embed.py chat_template.jinja processor_config.json
tokenizer.json tokenizer_config.json
backbone/ # frozen Qwen3-VL backbone (vocab-resized), safetensors
whisper_encoder.pt dasheng_encoder.pt
projector_weights.pt # projectors; LoRA already merged into the encoders
The .pt files are PyTorch pickles, and the custom modeling code requires
trust_remote_code=True. Load only from a source you trust.
@inproceedings{omniembedmini2026,
title = {Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation},
author = {TBD},
booktitle = {TBD},
year = {2026},
note = {Camera-ready in preparation}
}
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys omni-embed-mini for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (omni-embed-mini below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"omni-embed-mini","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.