Model reference · open weights

mmarco-mMiniLM-L12-H384

mmarco-mMiniLM-L12-H384 is an open-weight embedding model from cross-encoder, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings cross-encoder 1 variants 1.8M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What mmarco-mMiniLM-L12-H384 is

Cross-Encoder for multilingual MS Marco This model was trained on the MMARCO dataset. It is a machine translated version of MS MARCO using Google Translate. It was translated to 14 languages. In our experiments, we observed that it performs also well for other languages. As a base model, we used the multilingual MiniLMv2 model. The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. The training code is available here: SBERT.net Training MS Marco Usage with SentenceTransformers The usage becomes easy when you have SentenceTransformers installed. Then, you can use the pre-trained models like this: Usage with Transformers

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makercross-encoder
TypeEmbedding models
Parameters (lead)118M
Context514 tokens
Variants1
Runs withsentence-transformers
Based onnreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large
Released2022-06-01
Popularity1.8M downloads / month
Likes79
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
mmarco-mMiniLMv2-L12-H384-v1118MBF16~0.3 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys mmarco-mminilm-l12-h384 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (mmarco-mminilm-l12-h384 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"mmarco-mminilm-l12-h384","input":"text to embed"}'

Details

Languages, data & research

Languages

en ar zh nl fr de hi in it ja pt ru es vi

Trained / evaluated on

unicamp-dl/mmarco

Tags

sentence-transformers pytorch onnx safetensors openvino xlm-roberta text-classification transformers text-ranking en ar zh nl fr

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want mmarco-mMiniLM-L12-H384 on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms