Model reference · open weights
mMiniLMv2-L6-H384 is an open-weight embedding model from hotchpotch. mMiniLMv2-L6-H384 (FP32) weighs 214 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | hotchpotch |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 107M |
| Context | 514 tokens |
| Runs with | transformers |
| Released | 2024-03-04 |
| Popularity | 3k downloads / month |
| Weights | 214 MB (mMiniLMv2-L6-H384 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 214 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This model is a re-upload of Microsoft's Multilingual MiniLM v2 to HuggingFace under the MIT License, made more accessible for use through the HuggingFace transformers library.
The original pre-trained model is provided at the following URL:
The license for this model is based on the original license (found in the LICENSE file in the project's root directory), which is the MIT License.
All credits for this model go to the authors of Multilingual MiniLM v2 and the associated researchers and organizations. When using this model, please be sure to attribute the original authors.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.