Model reference · open weights
USER2-small is an open-weight embedding model from deepvk. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | deepvk |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 34M |
| Context | 8k tokens |
| Runs with | sentence-transformers |
| Based on | deepvk/RuModernBERT-small |
| Released | 2025-02-19 |
| Popularity | 7k downloads / month |
| Licence | Open weights |
About
USER2 is a new generation of the Universal Sentence Encoder for Russian, designed for sentence representation with long-context support of up to 8,192 tokens.
The models are built on top of the RuModernBERT encoders and are fine-tuned for retrieval and semantic tasks.
They also support Matryoshka Representation Learning (MRL) — a technique that enables reducing embedding size with minimal loss in representation quality.
This is a small model with 34 million parameters.
| Model | Size | Context Length | Hidden Dim | MRL Dims |
|---|---|---|---|---|
deepvk/USER2-small | 34M | 8192 | 384 | [32, 64, 128, 256, 384] |
deepvk/USER2-base | 149M | 8192 | 768 | [32, 64, 128, 256, 384, 512, 768] |
To evaluate the model, we measure quality on the MTEB-rus benchmark.
Additionally, to measure long-context retrieval, we run Russian subset of MultiLongDocRetrieval (MLDR) task.
MTEB-rus
| Model | Size | Hidden Dim | Context Length | MRL support | Mean(task) | Mean(taskType) | Classification | Clustering | MultiLabelClassification | PairClassification | Reranking | Retrieval | STS |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
USER-base | 124M | 768 | 512 | ❌ | 58.11 | 56.67 | 59.89 | 53.26 | 37.72 | 59.76 | 55.58 | 56.14 | 74.35 |
USER-bge-m3 | 359M | 1024 | 8192 | ❌ | 62.80 | 62.28 | 61.92 | 53.66 | 36.18 | 65.07 | 68.72 | 73.63 | 76.76 |
multilingual-e5-base | 278M | 768 | 512 | ❌ | 58.34 | 57.24 | 58.25 | 50.27 | 33.65 | 54.98 | 66.24 | 67.14 | 70.16 |
multilingual-e5-large-instruct | 560M | 1024 | 512 | ❌ | 65.00 | 63.36 | 66.28 | 63.13 | 41.15 | 63.89 | 64.35 | 68.23 | 76.48 |
jina-embeddings-v3 | 572M | 1024 | 8192 | ✅ | 63.45 | 60.93 | 65.24 | 60.90 | 39.24 | 59.22 | 53.86 | 71.99 | 76.04 |
ru-en-RoSBERTa | 404M | 1024 | 512 | ❌ | 61.71 | 60.40 | 62.56 | 56.06 | 38.88 | 60.79 | 63.89 | 66.52 | 74.13 |
USER2-small | 34M | 384 | 8192 | ✅ | 58.32 | 56.68 | 59.76 | 57.06 | 33.56 | 54.02 | 58.26 | 61.87 | 72.25 |
USER2-base | 149M | 768 | 8192 | ✅ | 61.12 | 59.59 | 61.67 | 59.22 | 36.61 | 56.39 | 62.06 | 66.90 | 74.28 |
MLDR-rus
| Model | Size | nDCG@10 ↑ |
|---|---|---|
USER-bge-m3 | 359M | 58.53 |
KaLM-v1.5 | 494M | 53.75 |
jina-embeddings-v3 | 572M | 49.67 |
E5-mistral-7b | 7.11B | 52.40 |
USER2-small | 34M | 51.69 |
USER2-base | 149M | 54.17 |
We compare only model with context length of 8192.
To evaluate MRL capabilities, we also use MTEB-rus, applying dimensionality cropping to the embeddings to match the selected size.
This model is trained similarly to Nomic Embed and expects task-specific prefixes to be added to the input. The choice of prefix depends on the specific task. We follow a few general guidelines when selecting a prefix:
However, we encourage users to experiment with different prefixes, as certain domains may benefit from specific ones.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("deepvk/USER2-small")
query_embeddings = model.encode(["Когда был спущен на воду первый миноносец «Спокойный»?"], prompt_name="search_query")
document_embeddings = model.encode(["Спокойный (эсминец)\nЗачислен в списки ВМФ СССР 19 августа 1952 года."], prompt_name="search_document")
similarities = model.similarity(query_embeddings, document_embeddings)
To truncate the embedding dimension, simply pass the new value to the model initialization:
model = SentenceTransformer("deepvk/USER2-small", truncate_dim=128)
This model was trained with dimensions [32, 64, 128, 256, 384], so it’s recommended to use one of these for best performance.
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
def mean_pooling(model_output, attention_mask):
token_embeddings = model_output[0]
input_mask_expanded = (
attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
)
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(
input_mask_expanded.sum(1), min=1e-9
)
queries = ["search_query: Когда был спущен на воду первый миноносец «Спокойный»?"]
documents = ["search_documenFrom the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys user2-small for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (user2-small below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"user2-small","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.