Model reference · open weights

Omni-Embed-Mini-onnx

Embeddings MBZUAI Embeddings 1 build Open weights 24 dl/mo

Omni-Embed-Mini-onnx is an open-weight embedding model from MBZUAI.

  • Omni-Embed-Mini-onnx is a 0.9B multimodal embedding model by MBZUAI designed for feature extraction in browser environments using WebGPU.
  • It embeds text, images, video, speech, general audio, and document pages into a single 1024-dimensional space.
  • The model is released under the Apache 2.0 license and is available in fp16 and q8 precisions, with the default fp16 bundle sized at 1.89 GB.

Summary of the MBZUAI/Omni-Embed-Mini-0.9B-onnx model card, 2026-10-03

What it is

Released byMBZUAI
Released2026-09-13

From the model card

What MBZUAI says about Omni-Embed-Mini-onnx

Read the model card

ONNX export of the Omni-Embed Mini 0.9B multimodal embedding model, built to run in a browser on WebGPU. One model embeds text, images, video, speech, general audio and document pages into a single 1024-dimensional space.

  • Project page: https://omniembed.cvmbzuai.com/
  • Live demo: https://demo-omniembed.cvmbzuai.com/
  • Paper: https://huggingface.co/papers/2610.02148
  • GitHub: https://github.com/k-m-irfan/Omni-Embed-Mini

What is in here

Two directories, one per precision, each a complete bundle with the same file names, so a client picks one by changing a base URL. fp16/ is the default at 1.89 GB. q8/ is 1.39 GB and about half the memory in use, for devices that cannot hold fp16; measured on a 6,000 item index it returns the same best result for 12 of 14 queries and the same 9.8 of the top 10.

Each is a six-graph bundle driven from JavaScript rather than a transformers.js architecture: the splice, the pooling and two of the vision tower's three inputs are computed by the caller. See manifest.json for the preconditions a caller cannot read off the graphs.

ComponentGraph
backbonebackbone/model.onnx, takes inputs_embeds
embedding tablebackbone/embed_tokens.onnx
visionvision_encoder/model.onnx, one image per call
audiowhisper_encoder/, dasheng_encoder/
projectorsprojectors/*.onnx

Every component was parity-checked against its PyTorch reference before export. The measured numbers are in */conversion_metadata.json and */parity_report.json.

Preprocessing is part of the model: a still image is squared to 224 with PIL bilinear and then taken to 256 with a Pillow-compatible bicubic, and substituting a browser canvas resize moves the final embedding by cos 0.734, so a client that does not reproduce that chain is a different model and its vectors must not be mixed into an index built with this one.

Citation

@misc{kurpath2026omniembedminibindingmodalitiesforgetting,
      title={Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation},
      author={Mohammed Irfan Kurpath and Jaseel Muhammad Kaithakkodan and Sahal Shaji Mullappilly and Ivan Laptev and Hisham Cholakkal},
      year={2026},
      eprint={2610.02148},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02148},
}

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms