Model reference · open weights
jina-embeddings-text-small-clustering is an open-weight embedding model from jinaai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | jinaai |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 596M |
| Context | 32k tokens |
| Runs with | llama.cpp |
| Based on | jinaai/jina-embeddings-v5-text-small |
| Released | 2026-02-10 |
| Popularity | 2k downloads / month |
| Licence | Commercial licence needed |
About
Elastic Inference Service | ArXiv | Release Note | Blog
jina-embeddings-v5-text-small-clustering is a compact, high-performance text embedding model designed for clustering.
It is part of the jina-embeddings-v5-text model family, which also includes jina-embeddings-v5-text-nano, a smaller model for more resource-constrained use cases.
Trained using a novel approach that combines distillation with task-specific contrastive losses, jina-embeddings-v5-text-small-clustering outperforms existing state-of-the-art models of similar size across diverse embedding benchmarks.
| Feature | Value |
|---|---|
| Parameters | 677M |
| Supported Tasks | clustering |
| Max Sequence Length | 32768 |
| Embedding Dimension | 1024 |
| Matryoshka Dimensions | 32, 64, 128, 256, 512, 768, 1024 |
| Pooling Strategy | Last-token pooling |
| Base Model | jinaai/jina-embeddings-v5-text-small |
For training details and evaluation results, see our technical report.
The following Python packages are required:
transformers>=5.1.0torch>=2.8.0peft>=0.15.2vllm>=0.15.1sentence-transformers interface, install this package as well.The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.
PUT _inference/text_embedding/jina-v5
{
"service": "elastic",
"service_settings": {
"model_id": "jina-embeddings-v5-text-small"
}
}
See the Elastic Inference Service documentation for setup details.
from sentence_transformers import SentenceTransformer
import torch
model = SentenceTransformer(
"jinaai/jina-embeddings-v5-text-small-clustering",
model_kwargs={"dtype": torch.bfloat16}, # Recommended for GPUs
config_kwargs={"_attn_implementation": "flash_attention_2"}, # Recommended but optional
)
# Optional: set truncate_dim in encode() to control embedding size
texts = [
"We propose a novel neural network architecture for image segmentation.",
"This paper analyzes the effects of monetary policy on inflation.",
"Our method achieves state-of-the-art results on object detection benchmarks.",
"We study the relationship between interest rates and housing prices.",
"A new attention mechanism is introduced for visual recognition tasks.",
]
# Encode texts
embeddings = model.encode(texts)
print(embeddings.shape)
# (5, 1024)
similarity = model.similarity(embeddings, embeddings)
print(similarity)
# tensor([[1.0000, 0.2983, 0.8631, 0.3098, 0.9106],
# [0.2983, 1.0000, 0.3257, 0.8041, 0.3201],
# [0.8631, 0.3257, 1.0000, 0.3263, 0.9007],
# [0.3098, 0.8041, 0.3263, 1.0000, 0.3122],
# [0.9106, 0.3201, 0.9007, 0.3122, 1.0000]])
from vllm import LLM#
from vllm.config.pooler import PoolerConfig
# Initialize model
name = "jinaai/jina-embeddings-v5-text-small-clustering"
model = LLM(
model=name,
dtype="float16",
runner="pooling",
pooler_config=PoolerConfig(seq_pooling_type="LAST", normalize=True)
)
# Create text prompts
document1 = "Overview of climate change impacts on coastal cities"
document1_prompt = f"Document: {document1}"
document2 = "The impacts of climate change on large cities"
document2_prompt = f"Document: {document2}"
# Encode all prompts
prompts = [document1_prompt, document2_prompt]
outputs = model.encode(prompts, pooling_task="embed")
embed_document1 = outputs[0].outputs.data
embed_document2 = outputs[1].outputs.data
docker run -p 8080:80 \
ghcr.io/huggingface/text-embeddings-inference:cpu-1.9 \
--model-id jinaai/jina-embeddings-v5-text-small-clustering \
--dtype float32 --pooling last-token
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 \
--model-id jinaai/jina-embeddings-v5-text-small-clustering \
--dtype float16 --pooling last-token
Alternatively, you can also run with
cargo, more information can be found in the Text Embeddings Inference documentation.
Send a request to /v1/embeddings to generate embeddings via the OpenAI Embeddings API:
curl -X POST http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{
"model": "jinaai/jina-embeddings-v5-text-small-clustering",
"input": [
"Document: The impacts of climate change on coastal cities are significant...",
]
}'
Or rather via the Text Embeddings Inference API specification instead, to prevent from manually formatting the inputs:
curl -X POST http://127.0.0.1:8080/embed \
-H "Content-Type: application/json" \
-d '{
"inputs": "Overview of climate change impacts on coastal cities",
"prompt_name": "document",
}'
After installing llam
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys jina-embeddings-text-small-clustering for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (jina-embeddings-text-small-clustering below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"jina-embeddings-text-small-clustering","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.