Model reference · open weights

mmbert-embed-32k-2d-matryoshka

Available as managed deployment Embeddings llm-semantic-router Embeddings 1 variants 34k dl/mo

mmbert-embed-32k-2d-matryoshka is an open-weight embedding model from llm-semantic-router. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byllm-semantic-router
TypeEmbedding models
TaskEmbeddings
Parameters (lead)307M
Context32k tokens
Runs withsentence-transformers
Based onllm-semantic-router/mmbert-32k-yarn
Released2026-01-26
Popularity34k downloads / month
LicenceOpen weights

About

What mmbert-embed-32k-2d-matryoshka is

A multilingual embedding model with 32K context window and 2D Matryoshka support for flexible efficiency-quality tradeoffs.

Read the full model card

Model Highlights

FeatureValue
Parameters307M
Context Length32,768 tokens
Languages1800+ (via Glot500)
Embedding Dim768 (supports 64-768 via Matryoshka)
ArchitectureModernBERT encoder with YaRN scaling

Key Results

MetricScore
MTEB Mean (24 tasks)61.4
STS Benchmark80.5 (exceeds Qwen3-0.6B's 76.17)
Dimension Retention99% @ 256d, 98% @ 64d
Layer Speedup3.3× @ 6L, 5.8× @ 3L
Latency vs BGE-M31.6-3.1× faster (FA2 advantage)

What is 2D Matryoshka?

This model supports two dimensions of flexibility:

  1. Dimension Reduction (Matryoshka): Truncate embeddings to smaller dimensions with minimal quality loss
  2. Layer Reduction (Adaptive): Use intermediate layer outputs for faster inference
ConfigQualitySpeedupStorage
22L, 768d100%1.0×100%
22L, 256d99%1.0×33%
22L, 64d98%1.0×8%
6L, 768d56%3.3×100%
6L, 256d56%3.3×33%

Usage

Basic Usage (Sentence Transformers)

from sentence_transformers import SentenceTransformer

# Load model
model = SentenceTransformer("llm-semantic-router/mmbert-embed-32k-2d-matryoshka")

# Encode sentences
sentences = [
    "This is a test sentence.",
    "这是一个测试句子。",
    "Dies ist ein Testsatz.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)  # (3, 768)

Matryoshka Dimension Reduction

import torch.nn.functional as F

# Encode with full dimensions
embeddings = model.encode(sentences, convert_to_tensor=True)

# Truncate to smaller dimension (e.g., 256)
embeddings_256d = embeddings[:, :256]
embeddings_256d = F.normalize(embeddings_256d, p=2, dim=1)

# Or truncate to 64 dimensions for maximum compression
embeddings_64d = embeddings[:, :64]
embeddings_64d = F.normalize(embeddings_64d, p=2, dim=1)

Long Context (up to 32K tokens)

# For long documents, set max_seq_length
model.max_seq_length = 8192  # or up to 32768

long_document = "..." * 10000  # Very long text
embedding = model.encode(long_document)

Layer Reduction (Advanced)

For latency-critical applications, you can extract embeddings from intermediate layers:

from transformers import AutoModel, AutoTokenizer
import torch
import torch.nn.functional as F

model = AutoModel.from_pretrained(
    "llm-semantic-router/mmbert-embed-32k-2d-matryoshka",
    trust_remote_code=True,
    output_hidden_states=True
)
tokenizer = AutoTokenizer.from_pretrained("llm-semantic-router/mmbert-embed-32k-2d-matryoshka")

# Encode
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

    # Use layer 6 for 3.3× speedup (56% quality)
    hidden = outputs.hidden_states[6]
    hidden = model.final_norm(hidden)

    # Mean pooling
    mask = inputs["attention_mask"].unsqueeze(-1).float()
    pooled = (hidden * mask).sum(1) / mask.sum(1)
    embeddings = F.normalize(pooled, p=2, dim=1)

Evaluation Results

MTEB Benchmark (24 tasks)

CategoryScore
STS (7 tasks)79.3
Classification (6)62.4
Pair Classification (2)76.2
Reranking (2)64.4
Clustering (4)36.9
Retrieval (3)38.2
Overall Mean61.4

STS Benchmark

ModelParametersSTS Score
Qwen3-Embed-0.6B600M76.17
mmBERT-Embed307M80.5
Qwen3-Embed-8B8B81.08

2D Matryoshka Quality Matrix (STS)

Layers768d256d64d
22L80.579.978.5
11L53.748.044.4
6L45.245.243.5
3L44.044.141.8

Long-Context Retrieval (4K tokens)

MetricScore
R@168.8%
R@1081.2%
MRR71.9%

Throughput (AMD MI300X)

LayersThroughputSpeedup
22L477/s1.0×
11L916/s1.9×
6L1573/s3.3×
3L2761/s5.8×

Latency Comparison vs BGE-M3 and Qwen3-Embedding-0.6B

mmBERT-Embed is significantly faster due to:

  1. Flash Attention 2 - BGE-M3 lacks FA2 (O(n) vs O(n²) attention)
  2. Encoder architecture - Qwen3 uses decoder with causal masking
  3. Smaller model - 307M vs 569M/600M params
Batch Size = 1
Seq LenmmBERT-EmbedQwen3-0.6BBGE-M3mmBERT Speedup
51217.6ms (57/s)20.7ms (48/s)10.8ms (93/s)0.6×
102418.6ms (54/s)21.2ms (47/s)16.3ms (61/s)0.9×
204819.5ms (51/s)24.1ms (42/s)31.1ms (32/s)1.6×
409621.3ms (47/s)43.5ms (23/s)60.5ms (17/s)2.8×
Batch Size = 8
Seq LenmmBERT-EmbedQwen3-0.6BBGE-M3mmBERT Speedup
51221.1ms (379/s)33.0ms (243/s)40.0ms (200/s)1.9×
102434.5ms (232/s)58.5ms (137/s)77.4ms (103/s)2.2×
204865.2ms (123/s)117.0ms (68/s)162.9ms (49/s)2.5×
4096130.7ms (61/s)254.9ms (31/s)411.3ms (19/s)3.1×

Key insight: The FA2 advantage grows with sequence length and batch size:

  • At short sequences (512), BGE-M3 is faster (no FA2 overhead)
  • At 2K+ tokens, mmBERT pulls ahead significantly
  • At 4K batch=8: mmBERT is 3.1× faster than BGE-M3

Benchmarked on AMD MI300X, bf16 precision.

Training

Data

Trained on BAAI/bge-m3-data (73GB, 279 JSONL files) with:

  • Multilingual triplets (query, positive, negative)
  • Diverse do

From the published model card. Full card on the HuggingFace links in the sidebar.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
STSSTS Benchmarkspearman80.500

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys mmbert-embed-32k-2d-matryoshka for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (mmbert-embed-32k-2d-matryoshka below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"mmbert-embed-32k-2d-matryoshka","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms