Model reference · open weights

ms-marco-MiniLM-L4

ms-marco-MiniLM-L4 is an open-weight embedding model from cross-encoder, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings cross-encoder 1 variants 6.4M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What ms-marco-MiniLM-L4 is

Cross-Encoder for MS Marco This model was trained on the MS Marco Passage Ranking task. The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. The training code is available here: SBERT.net Training MS Marco Usage with SentenceTransformers The usage is easy when you have SentenceTransformers installed. Then you can use the pre-trained models like this: Usage with Transformers Performance In the following table, we provide various pre-trained Cross-Encoders together with their performance on the TREC Deep Learning 2019 and the MS Marco Passage Reranking dataset. Note: Runtime was computed on a V100 GPU.

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makercross-encoder
TypeEmbedding models
Parameters (lead)19M
Context512 tokens
Variants1
Runs withsentence-transformers
Based oncross-encoder/ms-marco-MiniLM-L12-v2
Released2022-03-02
Popularity6.4M downloads / month
Likes28
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
ms-marco-MiniLM-L4-v219MBF16~0 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys ms-marco-minilm-l4 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (ms-marco-minilm-l4 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"ms-marco-minilm-l4","input":"text to embed"}'

Details

Languages, data & research

Languages

en

Trained / evaluated on

sentence-transformers/msmarco

Tags

sentence-transformers pytorch jax onnx safetensors openvino bert text-classification transformers text-ranking en dataset:sentence-transformers/msmarco text-embeddings-inference endpoints_compatible

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want ms-marco-MiniLM-L4 on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms