Model reference · open weights
MiniCPM-Embedding is an open-weight embedding model from openbmb. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | openbmb |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 2.7B |
| Context | 512 tokens |
| Runs with | transformers |
| Based on | openbmb/MiniCPM-2B-sft-bf16 |
| Released | 2024-09-04 |
| Popularity | 12k downloads / month |
| Licence | Unknown |
About
MiniCPM-Embedding 是面壁智能与清华大学自然语言处理实验室(THUNLP)、东北大学信息检索小组(NEUIR)共同开发的中英双语言文本嵌入模型,有如下特点:
MiniCPM-Embedding 基于 MiniCPM-2B-sft-bf16 训练,结构上采取双向注意力和 Weighted Mean Pooling [1]。采取多阶段训练方式,共使用包括开源数据、机造数据、闭源数据在内的约 600 万条训练数据。
欢迎关注 RAG 套件系列:
MiniCPM-Embedding is a bilingual & cross-lingual text embedding model developed by ModelBest Inc. , THUNLP and NEUIR , featuring:
MiniCPM-Embedding is trained based on MiniCPM-2B-sft-bf16 and incorporates bidirectional attention and Weighted Mean Pooling [1] in its architecture. The model underwent multi-stage training using approximately 6 million training examples, including open-source, synthetic, and proprietary data.
We also invite you to explore the RAG toolkit series:
[1] Muennighoff, N. (2022). Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904.
模型大小:2.4B
嵌入维度:2304
最大输入token数:512
Model Size: 2.4B
Embedding Dimension: 2304
Max Input Tokens: 512
本模型支持 query 侧指令,格式如下:
MiniCPM-Embedding supports query-side instructions in the following format:
Instruction: {{ instruction }} Query: {{ query }}
例如:
For example:
Instruction: 为这个医学问题检索相关回答。Query: 咽喉癌的成因是什么?
Instruction: Given a claim about climate change, retrieve documents that support or refute the claim. Query: However the warming trend is slower than most climate models have forecast.
也可以不提供指令,即采取如下格式:
MiniCPM-Embedding also works in instruction-free mode in the following format:
Query: {{ query }}
我们在 BEIR 与 C-MTEB/Retrieval 上测试时使用的指令见 instructions.json,其他测试不使用指令。文档侧直接输入文档原文。
When running evaluation on BEIR and C-MTEB/Retrieval, we use instructions in instructions.json. For other evaluations, we do not use instructions. On the document side, we directly use the bare document as the input.
transformers==4.37.2
from transformers import AutoModel, AutoTokenizer
import torch
import torch.nn.functional as F
model_name = "openbmb/MiniCPM-Embedding"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.float16).to("cuda")
# You can also use the following line to enable the Flash Attention 2 implementation
# model = AutoModel.from_pretrained(model_name, trust_remote_code=True, attn_implementation="flash_attention_2", torch_dtype=torch.float16).to("cuda")
model.eval()
# 由于在 `model.forward` 中缩放了最终隐层表示,此处的 mean pooling 实际上起到了 weighted mean pooling 的作用
# As we scale hidden states in `model.forward`, mean pooling here actually works as weighted mean pooling
def mean_pooling(hidden, attention_mask):
s = torch.sum(hidden * attention_mask.unsqueeze(-1).float(), dim=1)
d = attention_mask.sum(dim=1, keepdim=True).float()
reps = s / d
return reps
@torch.no_grad()
def encode(input_texts):
batch_dict = tokenizer(input_texts, max_length=512, padding=True, truncation=True, return_tensors='pt', return_attention_mask=True).to("cuda")
outputs = model(**batch_dict)
attention_mask = batch_dict["attention_mask"]
hidden = outputs.last_hidden_state
reps = mean_pooling(hidden, attention_mask)
embeddings = F.normalize(reps, p=2, dim=1).detach().cpu().numpy()
return embeddings
queries = ["中国的首都是哪里?"]
passages = ["beijing", "shanghai"]
INSTRUCTION = "Query: "
queries = [INSTRUCTION + query for query in queries]
embeddings_query = encode(queries)
embeddings_doc = encode(passages)
scores = (embeddings_query @ embeddings_doc.T)
print(scores.tolist()) # [[0.3535913825035095, 0.18596848845481873]]
import torch
from sentence_transformers import SentenceTransformer
model_name = "openbmb/MiniCPM-Embedding"
model = SentenceTransformer(model_name, trust_remote_code=True, model_kwargs={ "torch_dtype": torch.float16})
# You can also use the following line to enable the Flash Attention 2 implementation
# model = SentenceTransformer(model_name, trust_remote_code=True, attn_implementation="flash_attention_2", model_kwargs={ "torch_dtype": torch.float16})
queries = ["中国的首都是哪里?"]
passages = ["beijing", "shanghai"]
INSTRUCTION = "Query: "
embeddings_query = model.encode(queries, prompt=INSTRUCTION)
embeddings_doc = model.encode(passages)
scores = (embeddings_query @ embeddings_doc.T)
print(scores.tolist()) # [[0.35365450382232666, 0.18592746555805206]]
| 模型 Model | C-MTEB/Retrieval (NDCG@10) | BEIR (NDCG@10) |
|---|---|---|
| bge-large-zh-v1.5 | 70.46 | - |
| gte-large-zh | 72.49 | - |
| Zhihui_LLM_Embedding | 76.74 | |
| bge-large-en-v1.5 | - | 54.29 |
| gte-en-large-v1.5 | - | 57.91 |
| NV-Retriever-v1 | - | 60.9 |
| bge-en-icl |
From the published model card. Full card on the HuggingFace links in the sidebar.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Retrieval | MTEB ArguAna | ndcg_at_10 | 64.650 |
| Retrieval | MTEB CQADupstackRetrieval | ndcg_at_10 | 46.530 |
| Retrieval | MTEB ClimateFEVER | ndcg_at_10 | 35.550 |
| Retrieval | MTEB DBPedia | ndcg_at_10 | 47.820 |
| Retrieval | MTEB FEVER | ndcg_at_10 | 90.760 |
| Retrieval | MTEB FiQA2018 | ndcg_at_10 | 56.640 |
| Retrieval | MTEB HotpotQA | ndcg_at_10 | 78.110 |
| Retrieval | MTEB MSMARCO | ndcg_at_10 | 43.930 |
| Retrieval | MTEB NFCorpus | ndcg_at_10 | 39.770 |
| Retrieval | MTEB NQ | ndcg_at_10 | 69.290 |
| Retrieval | MTEB QuoraRetrieval | ndcg_at_10 | 89.970 |
| Retrieval | MTEB SCIDOCS | ndcg_at_10 | 22.380 |
| Retrieval | MTEB SciFact | ndcg_at_10 | 86.600 |
| Retrieval | MTEB TRECCOVID | ndcg_at_10 | 81.320 |
| Retrieval | MTEB Touche2020 | ndcg_at_10 | 25.080 |
| Retrieval | MTEB CmedqaRetrieval | ndcg_at_10 | 46.050 |
| Retrieval | MTEB CovidRetrieval | ndcg_at_10 | 92.010 |
| Retrieval | MTEB DuRetrieval | ndcg_at_10 | 90.980 |
| Retrieval | MTEB EcomRetrieval | ndcg_at_10 | 70.210 |
| Retrieval | MTEB MMarcoRetrieval | ndcg_at_10 | 85.550 |
| Retrieval | MTEB MedicalRetrieval | ndcg_at_10 | 63.910 |
| Retrieval | MTEB T2Retrieval | ndcg_at_10 | 87.330 |
| Retrieval | MTEB VideoRetrieval | ndcg_at_10 | 78.050 |
Using it via the API
Once AxForge deploys minicpm-embedding for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (minicpm-embedding below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"minicpm-embedding","input":"text to embed"}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.