Model reference · open weights

mmbert-embed-32k-2d-matryoshka

Embeddings vllm-sr Embeddings 1 build Open weights 31k dl/mo

mmbert-embed-32k-2d-matryoshka is an open-weight embedding model from vllm-sr. mmbert-embed-32k-2d-matryoshka (BF16) weighs 614 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • mmbert-embed-32k-2d-matryoshka is a 307M parameter multilingual embedding model designed for sentence similarity tasks.
  • It features a 32,768 token context window and supports 2D Matryoshka training, allowing for flexible tradeoffs between inference speed and quality through dimension and layer reduction.
  • The model is released under the Apache 2.0 license.

Summary of the vllm-sr/mmbert-embed-32k-2d-matryoshka model card, 2026-10-03

What it is

Released byvllm-sr
Released2026-01-26
Parameters307M
VRAM614 MB for the weights

What it runs on

Memory and cards for mmbert-embed-32k-2d-matryoshka (BF16)

614 MBweights, file size
1.1 GBruntime overhead
CardRunsMemory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

From the model card

What vllm-sr says about mmbert-embed-32k-2d-matryoshka

Read the model card

A multilingual embedding model with 32K context window and 2D Matryoshka support for flexible efficiency-quality tradeoffs.

Model Highlights

FeatureValue
Parameters307M
Context Length32,768 tokens
Languages1800+ (via Glot500)
Embedding Dim768 (supports 64-768 via Matryoshka)
ArchitectureModernBERT encoder with YaRN scaling

Key Results

MetricScore
MTEB Mean (24 tasks)61.4
STS Benchmark80.5 (exceeds Qwen3-0.6B's 76.17)
Dimension Retention99% @ 256d, 98% @ 64d
Layer Speedup3.3× @ 6L, 5.8× @ 3L
Latency vs BGE-M31.6-3.1× faster (FA2 advantage)

What is 2D Matryoshka?

This model supports two dimensions of flexibility:

  1. Dimension Reduction (Matryoshka): Truncate embeddings to smaller dimensions with minimal quality loss
  2. Layer Reduction (Adaptive): Use intermediate layer outputs for faster inference
ConfigQualitySpeedupStorage
22L, 768d100%1.0×100%
22L, 256d99%1.0×33%
22L, 64d98%1.0×8%
6L, 768d56%3.3×100%
6L, 256d56%3.3×33%

Usage

Basic Usage (Sentence Transformers)

from sentence_transformers import SentenceTransformer

# Load model
model = SentenceTransformer("vllm-sr/mmbert-embed-32k-2d-matryoshka")

# Encode sentences
sentences = [
    "This is a test sentence.",
    "这是一个测试句子。",
    "Dies ist ein Testsatz.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)  # (3, 768)

Matryoshka Dimension Reduction

import torch.nn.functional as F

# Encode with full dimensions
embeddings = model.encode(sentences, convert_to_tensor=True)

# Truncate to smaller dimension (e.g., 256)
embeddings_256d = embeddings[:, :256]
embeddings_256d = F.normalize(embeddings_256d, p=2, dim=1)

# Or truncate to 64 dimensions for maximum compression
embeddings_64d = embeddings[:, :64]
embeddings_64d = F.normalize(embeddings_64d, p=2, dim=1)

Long Context (up to 32K tokens)

# For long documents, set max_seq_length
model.max_seq_length = 8192  # or up to 32768

long_document = "..." * 10000  # Very long text
embedding = model.encode(long_document)

Layer Reduction (Advanced)

For latency-critical applications, you can extract embeddings from intermediate layers:

from transformers import AutoModel, AutoTokenizer
import torch
import torch.nn.functional as F

model = AutoModel.from_pretrained(
    "vllm-sr/mmbert-embed-32k-2d-matryoshka",
    trust_remote_code=True,
    output_hidden_states=True
)
tokenizer = AutoTokenizer.from_pretrained("vllm-sr/mmbert-embed-32k-2d-matryoshka")

# Encode
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

    # Use layer 6 for 3.3× speedup (56% quality)
    hidden = outputs.hidden_states[6]
    hidden = model.final_norm(hidden)

    # Mean pooling
    mask = inputs["attention_mask"].unsqueeze(-1).float()
    pooled = (hidden * mask).sum(1) / mask.sum(1)
    embeddings = F.normalize(pooled, p=2, dim=1)

Evaluation Results

MTEB Benchmark (24 tasks)

CategoryScore
STS (7 tasks)79.3
Classification (6)62.4
Pair Classification (2)76.2
Reranking (2)64.4
Clustering (4)36.9
Retrieval (3)38.2
Overall Mean61.4

STS Benchmark

ModelParametersSTS Score
Qwen3-Embed-0.6B600M76.17
mmBERT-Embed307M80.5
Qwen3-Embed-8B8B81.08

2D Matryoshka Quality Matrix (STS)

Layers768d256d64d
22L80.579.978.5
11L53.748.044.4
6L45.245.243.5
3L44.044.141.8

Long-Context Retrieval (4K tokens)

MetricScore
R@168.8%
R@1081.2%
MRR71.9%

Throughput (AMD MI300X)

LayersThroughputSpeedup
22L477/s1.0×
11L916/s1.9×
6L1573/s3.3×
3L2761/s5.8×

Latency Comparison vs BGE-M3 and Qwen3-Embedding-0.6B

mmBERT-Embed is significantly faster due to:

  1. Flash Attention 2 - BGE-M3 lacks FA2 (O(n) vs O(n²) attention)
  2. Encoder architecture - Qwen3 uses decoder with causal masking
  3. Smaller model - 307M vs 569M/600M params
Batch Size = 1
Seq LenmmBERT-EmbedQwen3-0.6BBGE-M3mmBERT Speedup
51217.6ms (57/s)20.7ms (48/s)10.8ms (93/s)0.6×
102418.6ms (54/s)21.2ms (47/s)16.3ms (61/s)0.9×
204819.5ms (51/s)24.1ms (42/s)31.1ms (32/s)1.6×
409621.3ms (47/s)43.5ms (23/s)60.5ms (17/s)2.8×
Batch Size = 8
Seq LenmmBERT-EmbedQwen3-0.6BBGE-M3mmBERT Speedup
51221.1ms (379/s)33.0ms (243/s)40.0ms (200/s)1.9×
102434.5ms (232/s)58.5ms (137/s)77.4ms (103/s)2.2×
204865.2ms (123/s)117.0ms (68/s)162.9ms (49/s)2.5×
4096130.7ms (61/s)254.9ms (31/s)411.3ms (19/s)3.1×

Key insight: The FA2 advantage grows with sequence length and batch size:

  • At short sequences (512), BGE-M3 is faster (no FA2 overhead)
  • At 2K+ tokens, mmBERT pulls ahead significantly
  • At 4K batch=8: mmBERT is 3.1× faster than BGE-M3

Benchmarked on AMD MI300X, bf16 precision.

Training

Data

Trained on BAAI/bge-m3-data (73GB, 279 JSONL files) with:

  • Multilingual triplets (query, positive, negative)
  • Diverse domains and languages

Configurati

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

Benchmarks

Reported results

As published on the model card: the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
STSSTS Benchmarkspearman80.500
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms