Model reference · open weights
mmarco-mMiniLMv2-L12-H384 is an open-weight embedding model from cross-encoder. mmarco-mMiniLMv2-L12-H384-v1 (FP32) weighs 235 MB; the smallest configuration that runs it is RTX 3060 12 GB.
mmarco-mMiniLMv2-L12-H384 is a cross-encoder model with 118M parameters designed for text-ranking tasks in information retrieval. It was trained on the MMARCO dataset, a machine-translated version of MS MARCO, and supports a context length of 514 tokens. The model is available under the Apache-2.0 license and operates in English, Arabic, Chinese, Dutch, French, German, Hindi, in, Italian, Japanese, Portuguese, and Russian.
Summary of the cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 model card, 2026-10-01
What it is
| Released by | cross-encoder |
|---|---|
| Type | Embedding models |
| Task | Reranker |
| Parameters (lead) | 118M |
| Context | 514 tokens |
| Runs with | sentence-transformers |
| Based on | nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large |
| Released | 2022-06-01 |
| Popularity | 1.8M downloads / month |
| Weights | 235 MB (mmarco-mMiniLMv2-L12-H384-v1 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 235 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This model was trained on the MMARCO dataset. It is a machine translated version of MS MARCO using Google Translate. It was translated to 14 languages. In our experiments, we observed that it performs also well for other languages.
As a base model, we used the multilingual MiniLMv2 model.
The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. The training code is available here: SBERT.net Training MS Marco
The usage becomes easy when you have SentenceTransformers installed. Then, you can use the pre-trained models like this:
from sentence_transformers import CrossEncoder
model = CrossEncoder('model_name')
scores = model.predict([('Query', 'Paragraph1'), ('Query', 'Paragraph2') , ('Query', 'Paragraph3')])
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model = AutoModelForSequenceClassification.from_pretrained('model_name')
tokenizer = AutoTokenizer.from_pretrained('model_name')
features = tokenizer(['How many people live in Berlin?', 'How many people live in Berlin?'], ['Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.', 'New York City is famous for the Metropolitan Museum of Art.'], padding=True, truncation=True, return_tensors="pt")
model.eval()
with torch.no_grad():
scores = model(**features).logits
print(scores)
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.