Model reference · open weights

jina-embeddings-text-nano-clustering

Available as managed deployment Licence fee Embeddings jinaai Embeddings 1 variants 1k dl/mo

jina-embeddings-text-nano-clustering is an open-weight embedding model from jinaai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byjinaai
TypeEmbedding models
TaskEmbeddings
Parameters (lead)212M
Context8k tokens
Runs withllama.cpp
Based onjinaai/jina-embeddings-v5-text-nano
Released2026-02-10
Popularity1k downloads / month
LicenceCommercial licence needed

About

What jina-embeddings-text-nano-clustering is

jina-embeddings-v5-text: Task-Targeted Embedding Distillation

Elastic Inference Service | ArXiv | Release Note | Blog

Model Overview

jina-embeddings-v5-text-nano-clustering is a compact, high-performance text embedding model designed for clustering.

Read the full model card

It is part of the jina-embeddings-v5-text model family, which also includes jina-embeddings-v5-text-small, for better performance at a bigger size.

Trained using a novel approach that combines distillation with task-specific contrastive losses, jina-embeddings-v5-text-nano-clustering outperforms existing state-of-the-art models of similar size across diverse embedding benchmarks.

FeatureValue
Parameters239M
Supported Tasksclustering
Max Sequence Length8192
Embedding Dimension768
Matryoshka Dimensions32, 64, 128, 256, 512, 768
Pooling StrategyLast-token pooling
Base Modeljinaai/jina-embeddings-v5-text-nano

Training and Evaluation

For training details and evaluation results, see our technical report.

Usage

The following Python packages are required:

  • transformers>=5.1.0
  • torch>=2.8.0
  • peft>=0.15.2
  • vllm==0.15.1

Optional / Recommended

  • flash-attention: Installing flash-attention is recommended for improved inference speed and efficiency, but not mandatory.
  • sentence-transformers: If you want to use the model via the sentence-transformers interface, install this package as well.

The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.

PUT _inference/text_embedding/jina-v5
{
  "service": "elastic",
  "service_settings": {
    "model_id": "jina-embeddings-v5-text-nano"
  }
}

See the Elastic Inference Service documentation for setup details.

from sentence_transformers import SentenceTransformer
import torch

model = SentenceTransformer(
    "jinaai/jina-embeddings-v5-text-nano-clustering",
    trust_remote_code=True,
    model_kwargs={"dtype": torch.bfloat16},  # Recommended for GPUs
    config_kwargs={"_attn_implementation": "flash_attention_2"},  # Recommended but optional
)
# Optional: set truncate_dim in encode() to control embedding size

texts = [
    "We propose a novel neural network architecture for image segmentation.",
    "This paper analyzes the effects of monetary policy on inflation.",
    "Our method achieves state-of-the-art results on object detection benchmarks.",
    "We study the relationship between interest rates and housing prices.",
    "A new attention mechanism is introduced for visual recognition tasks.",
]

# Encode texts
embeddings = model.encode(texts)
print(embeddings.shape)
# (5, 768)

similarity = model.similarity(embeddings, embeddings)
print(similarity)
# tensor([[1.0000, 0.2933, 0.9304, 0.2928, 0.8635],
#         [0.2933, 1.0000, 0.3062, 0.8083, 0.3035],
#         [0.9304, 0.3062, 1.0000, 0.2943, 0.8651],
#         [0.2928, 0.8083, 0.2943, 1.0000, 0.2827],
#         [0.8635, 0.3035, 0.8651, 0.2827, 1.0000]])
from vllm import LLM
from vllm.config.pooler import PoolerConfig

# Initialize model
name = "jinaai/jina-embeddings-v5-text-nano-clustering"
model = LLM(
    model=name,
    dtype="float16",
    runner="pooling",
    trust_remote_code=True,
    pooler_config=PoolerConfig(seq_pooling_type="LAST", normalize=True)
)

# Create text prompts
query = "Overview of climate change impacts on coastal cities"
query_prompt = f"Query: {query}"

document = "The impacts of climate change on coastal cities are significant.."
document_prompt = f"Document: {document}"

# Encode all prompts
prompts = [query_prompt, document_prompt]
outputs = model.encode(prompts, pooling_task="embed")

embed_query = outputs[0].outputs.data
embed_document = outputs[1].outputs.data

Since our nano model is based on jinaai/jina-embeddings-v5-text-nano, which is not yet supported by llama.cpp, we provide our own branch of llama.cpp, which implements the necessary changes to support it for now.

To start the OpenAI API compatible HTTP server, run with the respective model version:

llama-server \
  -hf jinaai/jina-embeddings-v5-text-nano-clustering:F16 \
  --embedding \
  --pooling last \
  --batch-size 8192 \
  --ubatch-size 8192 \
  --ctx-size 8192

Client:

curl -X POST "http://127.0.0.1:8080/v1/embeddings" \
  -H "Content-Type: application/json" \
  -d '{
    "input": [
      "Document: A beautiful sunset over the beach",
      "Document: Un beau coucher de soleil sur la plage",
      "Document: 海滩上美丽的日落",
      "Document: 浜辺に沈む美しい夕日",
      "Document: Golden sunlight melts into the horizon, painting waves in warm amber and rose, while the sky whispers goodnight to the quiet, endless sea."
    ]
  }'

Note: For the clustering variant, always add Document: prefix in front of your input as shown above.

You can run the ONNX-optimized version of the model locally using Hugging Face's optimum library. Make sure you have the required dependencies installed (e.g., pip install optimum[onnxruntime] transformers torch):

from optimum.onnxruntime import ORTModelForFeatureExtraction
from transformers import AutoTokenizer
import torch

model_id = "jinaai/jina-embeddings-v5-text-nano-clustering"

# 1. Load tokenizer and ONNX model
# We specify the 

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys jina-embeddings-text-nano-clustering for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (jina-embeddings-text-nano-clustering below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"jina-embeddings-text-nano-clustering","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms