Model reference · open weights
multi-modal-embed-small is an open-weight embedding model from llm-semantic-router. multi-modal-embed-small (FP32) weighs 675 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the llm-semantic-router/multi-modal-embed-small model card, 2026-10-01
What it is
| Released by | llm-semantic-router |
|---|---|
| Released | 2026-02-05 |
| Parameters | 338M |
| VRAM | 675 MB for the weights |
What it runs on
| Card | Runs | Memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
From the model card
A compact multimodal embedding model that unifies text, image, and audio representations in a shared semantic space. Part of the MoM (Mixture of Models) family.
multi-modal-embed-small is a lightweight multimodal encoder (~120M parameters) supporting:
| Feature | Description |
|---|---|
| Embedding Dimension | 384 (supports MRL truncation to 32, 64, 128, 256) |
| Image Resolution | 512×512 |
| Audio Input | Up to 30s, 16kHz (Whisper Mel spectrogram) |
| Modalities | Text, Image, Audio, Multimodal fusion |
| 2DMSE Support | Early exit at any encoder layer |
| Languages | English |
pip install torch transformers pillow safetensors
Two checkpoint formats are available:
model.pt (932 MB) - PyTorch formatmodel.safetensors (1.35 GB) - SafeTensors formatimport torch
import torch.nn as nn
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer, SiglipModel, SiglipProcessor, WhisperModel, WhisperFeatureExtractor
from huggingface_hub import hf_hub_download
class MultiModalEmbedder(nn.Module):
"""Standalone multimodal embedder - no external dependencies."""
def __init__(self):
super().__init__()
# Text encoder (384d, no projection needed)
self.text_tokenizer = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
self.text_encoder = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
# Image encoder (768d -> 384d projection)
self.image_processor = SiglipProcessor.from_pretrained("google/siglip-base-patch16-512")
self.image_encoder = SiglipModel.from_pretrained("google/siglip-base-patch16-512").vision_model
self.image_proj = nn.Linear(768, 384)
# Audio encoder (384d, no projection needed)
self.audio_processor = WhisperFeatureExtractor.from_pretrained("openai/whisper-tiny")
self.audio_encoder = WhisperModel.from_pretrained("openai/whisper-tiny").encoder
def encode_text(self, texts):
if isinstance(texts, str):
texts = [texts]
inputs = self.text_tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
inputs = {k: v.to(next(self.parameters()).device) for k, v in inputs.items()}
outputs = self.text_encoder(**inputs)
embeddings = outputs.last_hidden_state.mean(dim=1) # Mean pooling
return F.normalize(embeddings, p=2, dim=-1)
def encode_image(self, images):
inputs = self.image_processor(images=images, return_tensors="pt")
inputs = {k: v.to(next(self.parameters()).device) for k, v in inputs.items()}
outputs = self.image_encoder(**inputs)
embeddings = outputs.pooler_output
embeddings = self.image_proj(embeddings) # 768 -> 384
return F.normalize(embeddings, p=2, dim=-1)
def encode_audio(self, waveform):
# waveform: numpy array or tensor at 16kHz
if isinstance(waveform, torch.Tensor):
waveform = waveform.squeeze().numpy()
inputs = self.audio_processor(waveform, sampling_rate=16000, return_tensors="pt")
inputs = {k: v.to(next(self.parameters()).device) for k, v in inputs.items()}
outputs = self.audio_encoder(**inputs)
embeddings = outputs.last_hidden_state.mean(dim=1) # Mean pooling
return F.normalize(embeddings, p=2, dim=-1)
# Load model
model = MultiModalEmbedder()
# Download and load trained weights
checkpoint_path = hf_hub_download(
repo_id="llm-semantic-router/multi-modal-embed-small",
filename="model.pt"
)
state_dict = torch.load(checkpoint_path, map_location="cpu", weights_only=False)
# Load text encoder weights
model.text_encoder.load_state_dict({
k.replace("text_encoder.encoder.", ""): v
for k, v in state_dict.items()
if k.startswith("text_encoder.encoder.")
})
# Load image encoder and projection weights
model.image_encoder.load_state_dict({
k.replace("image_encoder.vision_encoder.", ""): v
for k, v in state_dict.items()
if k.startswith("image_encoder.vision_encoder.")
})
model.image_proj.load_state_dict({
k.replace("image_encoder.projection.", ""): v
for k, v in state_dict.items()
if k.startswith("image_encoder.projection.")
})
# Load audio encoder weights
model.audio_encoder.load_state_dict({
k.replace("audio_encoder.encoder.", ""): v
for k, v in state_dict.items()
if k.startswith("audio_encoder.encoder.")
})
model.eval()
print("Model loaded successfully!")
import torch.nn.functional as F
# Single text
text_embedding = model.encode_text("A photo of a cat") # Shape: [1, 384]
# Batch of texts
texts = ["A fluffy orange cat", "A golden retriever dog", "A red sports car"]
text_embeddings = model.encode_text(texts) # Shape: [3, 384]
# Compute similarity
similarities = F.cosine_similarity(text_embeddings[0:1], text_embeddings[1:], dim=-1)
print(f"Cat vs Dog: {similarities[0]:.3f}")
print(f"Cat vs Car: {similarities[1]:.3f}")
from PIL import Image
import requests
from io import BytesIO
# Load image
url = "https://upload.wikimedia.org/wikipedia/commons/thumb/3/3a/Cat03.jpg/1200px-Cat03.jpg"
image = Image.open(BytesIO(requests.get(url).content)).convert('RGB')
# Get embedding
image_embedding = model.encode_image(image) # Shape: [1, 384]
import torchaudio
# Load audio (16kHz)
waveform, sQuoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.
How it works
Benchmarks
As published on the model card: the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| image-text-retrieval | COCO | Image-to-Text R@1 | 41.880 |
| image-text-retrieval | COCO | Image-to-Text R@5 | 71.640 |
| image-text-retrieval | COCO | Image-to-Text R@10 | 82.160 |
| audio-text-retrieval | LibriSpeech | Audio-to-Text R@1 | 36.380 |
| audio-text-retrieval | LibriSpeech | Audio-to-Text R@5 | 68.220 |
| audio-text-retrieval | LibriSpeech | Audio-to-Text R@10 | 79.520 |