Model reference · open weights

mmlw-e5-large

Available as managed deployment Embeddings sdadas · community Embeddings 1 variants 567 dl/mo

mmlw-e5-large is an open-weight embedding model from sdadas. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bysdadas
TypeEmbedding models
TaskEmbeddings
Parameters (lead)560M
Context514 tokens
Runs withsentence-transformers
Released2023-11-17
Popularity567 downloads / month
LicenceOpen weights

About

What mmlw-e5-large is

MMLW (muszę mieć lepszą wiadomość) are neural text encoders for Polish. This is a distilled model that can be used to generate embeddings applicable to many tasks such as semantic similarity, clustering, information retrieval. The model can also serve as a base for further fine-tuning. It transforms texts to 1024 dimensional vectors. The model was initialized with multilingual E5 checkpoint, and then trained with multilingual knowledge distillation method on a diverse corpus of 60 million Polish-English text pairs. We utilised English FlagEmbeddings (BGE) as teacher models for distillation.

Read the full model card

Usage (Sentence-Transformers)

⚠️ Our embedding models require the use of specific prefixes and suffixes when encoding texts. For this model, queries should be prefixed with "query: " and passages with "passage: " ⚠️

You can use the model like this with sentence-transformers:

from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

query_prefix = "query: "
answer_prefix = "passage: "
queries = [query_prefix + "Jak dożyć 100 lat?"]
answers = [
    answer_prefix + "Trzeba zdrowo się odżywiać i uprawiać sport.",
    answer_prefix + "Trzeba pić alkohol, imprezować i jeździć szybkimi autami.",
    answer_prefix + "Gdy trwała kampania politycy zapewniali, że rozprawią się z zakazem niedzielnego handlu."
]
model = SentenceTransformer("sdadas/mmlw-e5-large")
queries_emb = model.encode(queries, convert_to_tensor=True, show_progress_bar=False)
answers_emb = model.encode(answers, convert_to_tensor=True, show_progress_bar=False)

best_answer = cos_sim(queries_emb, answers_emb).argmax().item()
print(answers[best_answer])
# Trzeba zdrowo się odżywiać i uprawiać sport.

Evaluation Results

  • The model achieves an Average Score of 61.17 on the Polish Massive Text Embedding Benchmark (MTEB). See MTEB Leaderboard for detailed results.
  • The model achieves NDCG@10 of 56.09 on the Polish Information Retrieval Benchmark. See PIRB Leaderboard for detailed results.

Acknowledgements

This model was trained with the A100 GPU cluster support delivered by the Gdansk University of Technology within the TASK center initiative.

Citation

@inproceedings{dadas2024pirb,
  title={PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods},
  author={Dadas, Slawomir and Pere{\l}kiewicz, Micha{\l} and Po{\'s}wiata, Rafa{\l}},
  booktitle={Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)},
  pages={12761--12774},
  year={2024}
}

From the published model card. Full card on the HuggingFace links in the sidebar.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
ClusteringMTEB 8TagsClusteringv_measure30.624
ClassificationMTEB AllegroReviewsaccuracy37.684
ClassificationMTEB AllegroReviewsf134.192
RetrievalMTEB ArguAna-PLmap_at_138.407
RetrievalMTEB ArguAna-PLmap_at_1055.147
RetrievalMTEB ArguAna-PLmap_at_10055.757
RetrievalMTEB ArguAna-PLmap_at_100055.761
RetrievalMTEB ArguAna-PLmap_at_351.268
RetrievalMTEB ArguAna-PLmap_at_553.697
RetrievalMTEB ArguAna-PLmrr_at_140.043
RetrievalMTEB ArguAna-PLmrr_at_1055.841
RetrievalMTEB ArguAna-PLmrr_at_10056.459
RetrievalMTEB ArguAna-PLmrr_at_100056.463
RetrievalMTEB ArguAna-PLmrr_at_352.074
RetrievalMTEB ArguAna-PLmrr_at_554.365
RetrievalMTEB ArguAna-PLndcg_at_138.407
RetrievalMTEB ArguAna-PLndcg_at_1063.248
RetrievalMTEB ArguAna-PLndcg_at_10065.717
RetrievalMTEB ArguAna-PLndcg_at_100065.790
RetrievalMTEB ArguAna-PLndcg_at_355.404
RetrievalMTEB ArguAna-PLndcg_at_559.760
RetrievalMTEB ArguAna-PLprecision_at_138.407
RetrievalMTEB ArguAna-PLprecision_at_108.862
RetrievalMTEB ArguAna-PLprecision_at_1000.991

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys mmlw-e5-large for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (mmlw-e5-large below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"mmlw-e5-large","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms