Model reference · open weights

msmarco-MiniLM-L6-en-de

Available as managed deployment Embeddings cross-encoder Reranker 1 variants 8k dl/mo

msmarco-MiniLM-L6-en-de is an open-weight embedding model from cross-encoder. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Makercross-encoder
TypeEmbedding models
TaskReranker
Parameters (lead)107M
Context512 tokens
Runs withsentence-transformers
Based onmicrosoft/Multilingual-MiniLM-L12-H384
Released2022-03-02
Popularity8k downloads / month
LicenceOpen weights

About

What msmarco-MiniLM-L6-en-de is

This is a cross-lingual Cross-Encoder model for EN-DE that can be used for passage re-ranking. It was trained on the MS Marco Passage Ranking task.

The model can be used for Information Retrieval: See SBERT.net Retrieve & Re-rank.

The training code is available in this repository, see train_script.py.

Usage with SentenceTransformers

When you have SentenceTransformers installed, you can use the model like this:

from sentence_transformers import CrossEncoder

model = CrossEncoder('model_name', max_length=512)

query = 'How many people live in Berlin?'
docs = ['Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.', 'New York City is famous for the Metropolitan Museum of Art.']
pairs = [(query, doc) for doc in docs]
scores = model.predict(pairs)

Usage with Transformers

With the transformers library, you can use the model like this:

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model = AutoModelForSequenceClassification.from_pretrained('model_name')
tokenizer = AutoTokenizer.from_pretrained('model_name')

features = tokenizer(['How many people live in Berlin?', 'How many people live in Berlin?'], ['Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.', 'New York City is famous for the Metropolitan Museum of Art.'],  padding=True, truncation=True, return_tensors="pt")

model.eval()
with torch.no_grad():
    scores = model(**features).logits
    print(scores)

Performance

The performance was evaluated on three datasets:

  • TREC-DL19 EN-EN: The original TREC 2019 Deep Learning Track: Given an English query and 1000 documents (retrieved by BM25 lexical search), rank documents with according to their relevance. We compute NDCG@10. BM25 achieves a score of 45.46, a perfect re-ranker can achieve a score of 95.47.
  • TREC-DL19 DE-EN: The English queries of TREC-DL19 have been translated by a German native speaker to German. We rank the German queries versus the English passages from the original TREC-DL19 setup. We compute NDCG@10.
  • GermanDPR DE-DE: The GermanDPR dataset provides German queries and German passages from Wikipedia. We indexed the 2.8 Million paragraphs from German Wikipedia and retrieved for each query the top 100 most relevant passages using BM25 lexical search with Elasticsearch. We compute MRR@10. BM25 achieves a score of 35.85, a perfect re-ranker can achieve a score of 76.27.

We also check the performance of bi-encoders using the same evaluation: The retrieved documents from BM25 lexical search are re-ranked using query & passage embeddings with cosine-similarity. Bi-Encoders can also be used for end-to-end semantic search.

Model-NameTREC-DL19 EN-ENTREC-DL19 DE-ENGermanDPR DE-DEDocs / Sec
BM2545.46-35.85-
Cross-Encoder Re-Rankers
cross-encoder/msmarco-MiniLM-L6-en-de-v172.4365.5346.771600
cross-encoder/msmarco-MiniLM-L12-en-de-v172.9466.0749.91900
svalabs/cross-electra-ms-marco-german-uncased (DE only)--53.67260
deepset/gbert-base-germandpr-reranking (DE only)--53.59260
Bi-Encoders (re-ranking)
sentence-transformers/msmarco-distilbert-multilingual-en-de-v2-tmp-lng-aligned63.3858.2837.88940
sentence-transformers/msmarco-distilbert-multilingual-en-de-v2-tmp-trained-scratch65.5158.6938.32940
svalabs/bi-electra-ms-marco-german-uncased (DE only)--34.31450
deepset/gbert-base-germandpr-question_encoder (DE only)--42.55450

Note: Docs / Sec gives the number of (query, document) pairs we can re-rank within a second on a V100 GPU.

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys msmarco-minilm-l6-en-de for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (msmarco-minilm-l6-en-de below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"msmarco-minilm-l6-en-de","input":"text to embed"}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms