Model reference · open weights
Mursit-Large-TR-Retrieval is an open-weight embedding model from newmindai. Mursit-Large-TR-Retrieval (FP32) weighs 807 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Mursit-Large-TR-Retrieval is a 403M parameter embedding model developed by newmindai for sentence-similarity tasks, specifically optimized for Turkish legal domain retrieval. The model supports a context length of 2048 tokens and operates in Turkish and English. It is distributed under the Apache-2.0 licence.
Summary of the newmindai/Mursit-Large-TR-Retrieval model card, 2026-10-01
What it is
| Released by | newmindai |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 404M |
| Context | 2,048 tokens |
| Runs with | sentence-transformers |
| Based on | newmindai/Mursit-Large |
| Released | 2026-01-16 |
| Popularity | 9k downloads / month |
| Weights | 807 MB (Mursit-Large-TR-Retrieval (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 807 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Mursit-Large-TR-Retrieval is a large-scale Turkish embedding model pre-trained entirely from scratch on Turkish-dominant corpora and fine-tuned for retrieval tasks. The model is based on ModernBERT-large architecture (403M parameters) and optimized specifically for Turkish legal domain applications. This model achieves strong performance on Turkish retrieval benchmarks with 56.87 MTEB Score and 46.56 Legal Score, ranking among the top Turkish embedding models.
Key Features:
Model Type: Embedding Parameters: 403M Base Model: newmindai/Mursit-Large Architecture: ModernBERT-large Embedding Dimension: 1,024 Max Sequence Length: 2,048 tokens
The model is based on ModernBERT-large architecture:
Pre-training:
Post-training for Embeddings:
The following visualization shows the model's performance compared to other Turkish language models:
Model Performance Comparison: Legal Score vs. MTEB Score. Embedding models (green triangles) show superior performance compared to MLM models. Mursit-Large-TR-Retrieval achieves strong performance with 56.87 MTEB Score and 46.56 Legal Score, ranking among the top Turkish embedding models.
This model was evaluated on the comprehensive MTEB-Turkish benchmark, which includes 17 tasks across 5 task types. The benchmark evaluates models on general Turkish NLP tasks as well as domain-specific legal retrieval tasks.
The following table presents comprehensive evaluation results across all models evaluated on the MTEB-Turkish benchmark. This model's results are highlighted in italics.
| Model | MTEB | Legal | Cls. | Clus. | Pair | Ret. | STS | Cont. | Reg. | Case | Params | Type |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| embeddinggemma-300m | 65.42 | 50.63 | 77.74 | 45.05 | 80.02 | 55.06 | 69.22 | 83.97 | 39.56 | 28.38 | 307M | Emb. |
| bge-m3 | 62.87 | 51.16 | 75.35 | 35.86 | 78.88 | 54.42 | 69.83 | 86.08 | 38.09 | 29.3 | 567M | Emb. |
| Mursit-Embed-Qwen3-1.7B-TR | 56.84 | 34.76 | 68.46 | 42.22 | 59.67 | 50.1 | 63.77 | 70.22 | 17.94 | 16.11 | 1.7B | CLM-E. |
| Mursit-Large-TR-Retrieval | 56.87 | 46.56 | 67.72 | 41.15 | 59.78 | 51.69 | 64.01 | 81.78 | 32.67 | 25.24 | 403M | Emb. |
| Mursit-Base-TR-Retrieval | 55.86 | 47.52 | 66.25 | 39.75 | 61.31 | 50.07 | 61.9 | 80.4 | 34.1 | 28.07 | 155M | Emb. |
| Mursit-Embed-Qwen3-4B-TR | 53.65 | 37.0 | 67.29 | 36.68 | 58.36 | 51.12 | 54.77 | 69.25 | 24.21 | 17.56 | 4B | CLM-E. |
| ------- | ------ | ------- | ------ | ------ | ------ | ------ | ----- | ------- | ------ | ------ | -------- | ------ |
| bert-base-turkish-uncased | 46.23 | 24.94 | 68.05 | 33.81 | 60.44 | 32.01 | 36.85 | 52.47 | 12.05 | 10.29 | 110M | MLM |
| turkish-large-bert-cased | 45.3 | 19.12 | 67.43 | 34.24 | 60. |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.