Model reference · open weights

jina-embeddings-text-small-text-matching

Available as managed deployment Licence fee Embeddings jinaai Embeddings 1 variants 11k dl/mo

jina-embeddings-text-small-text-matching is an open-weight embedding model from jinaai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Makerjinaai
TypeEmbedding models
TaskEmbeddings
Parameters (lead)596M
Context32k tokens
Runs withllama.cpp
Based onjinaai/jina-embeddings-v5-text-small
Released2026-02-10
Popularity11k downloads / month
LicenceCommercial licence needed

About

What jina-embeddings-text-small-text-matching is

jina-embeddings-v5-text-small-text-matching: Text-Matching-Targeted Embedding Distillation

Elastic Inference Service | ArXiv | Release Note | Blog

Model Overview

jina-embeddings-v5-text-small-text-matching is a compact, high-performance text embedding model designed for text-matching.

It is part of the jina-embeddings-v5-text model family, which also includes jina-embeddings-v5-text-nano, a smaller model for more resource-constrained use cases.

Trained using a novel approach that combines distillation with task-specific contrastive losses, jina-embeddings-v5-text-small-text-matching outperforms existing state-of-the-art models of similar size across diverse embedding benchmarks.

FeatureValue
Parameters677M
Supported Taskstext-matching
Max Sequence Length32768
Embedding Dimension1024
Matryoshka Dimensions32, 64, 128, 256, 512, 768, 1024
Pooling StrategyLast-token pooling
Base Modeljinaai/jina-embeddings-v5-text-small

Training and Evaluation

For training details and evaluation results, see our technical report.

Usage

The following Python packages are required:

  • transformers>=5.1.0
  • torch>=2.8.0
  • peft>=0.15.2
  • vllm>=0.15.1

Optional / Recommended

  • flash-attention: Installing flash-attention is recommended for improved inference speed and efficiency, but not mandatory.
  • sentence-transformers: If you want to use the model via the sentence-transformers interface, install this package as well.

The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.

PUT _inference/text_embedding/jina-v5
{
  "service": "elastic",
  "service_settings": {
    "model_id": "jina-embeddings-v5-text-small"
  }
}

See the Elastic Inference Service documentation for setup details.

from sentence_transformers import SentenceTransformer
import torch

model = SentenceTransformer(
    "jinaai/jina-embeddings-v5-text-small-text-matching",
    model_kwargs={"dtype": torch.bfloat16},  # Recommended for GPUs
    config_kwargs={"_attn_implementation": "flash_attention_2"},  # Recommended but optional
)
# Optional: set truncate_dim in encode() to control embedding size

texts = [
    "A beautiful sunset over the beach",  # English
    "غروب جميل على الشاطئ",  # Arabic
    "海滩上美丽的日落",  # Chinese
    "Un beau coucher de soleil sur la plage",  # French
    "Ein wunderschöner Sonnenuntergang am Strand",  # German
    "Ένα όμορφο ηλιοβασίλεμα πάνω από την παραλία",  # Greek
    "समुद्र तट पर एक खूबसूरत सूर्यास्त",  # Hindi
    "Un bellissimo tramonto sulla spiaggia",  # Italian
    "浜辺に沈む美しい夕日",  # Japanese
    "해변 위로 아름다운 일몰",  # Korean
]

# Encode texts
embeddings = model.encode(texts)
print(embeddings.shape)
# (10, 1024)

similarity = model.similarity(embeddings[0], embeddings[1:])
print(similarity)
# tensor([[0.7833, 0.8926, 0.9333, 0.9421, 0.7588, 0.9068, 0.9301, 0.8521, 0.8768]])
from vllm import LLM
from vllm.config.pooler import PoolerConfig

# Initialize model
name = "jinaai/jina-embeddings-v5-text-small-text-matching"
model = LLM(
    model=name,
    dtype="float16",
    runner="pooling",
    pooler_config=PoolerConfig(seq_pooling_type="LAST", normalize=True)
)

# Create text prompts
document1 = "Overview of climate change impacts on coastal cities"
document1_prompt = f"Document: {document1}"

document2 = "The impacts of climate change on large cities"
document2_prompt = f"Document: {document2}"

# Encode all prompts
prompts = [document1_prompt, document2_prompt]
outputs = model.encode(prompts, pooling_task="embed")

embed_document1 = outputs[0].outputs.data
embed_document2 = outputs[1].outputs.data
  • Via Docker on CPU:
    docker run -p 8080:80 \
      ghcr.io/huggingface/text-embeddings-inference:cpu-1.9 \
      --model-id jinaai/jina-embeddings-v5-text-small-text-matching \
      --dtype float32 --pooling last-token
    
  • Via Docker on NVIDIA GPU (Turing, Ampere, Ada Lovelace, Hopper or Blackwell):
    docker run --gpus all --shm-size 1g -p 8080:80 \
      ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 \
      --model-id jinaai/jina-embeddings-v5-text-small-text-matching \
      --dtype float16 --pooling last-token
    

Alternatively, you can also run with cargo, more information can be found in the Text Embeddings Inference documentation.

Send a request to /v1/embeddings to generate embeddings via the OpenAI Embeddings API:

curl -X POST http://127.0.0.1:8080/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jinaai/jina-embeddings-v5-text-small-text-matching",
    "input": [
      "Document: The impacts of climate change on coastal cities are significant...",
    ]
  }'

Or rather via the Text Embeddings Inference API specification instead, to prevent from manually formatting the inputs:

curl -X POST http://127.0.0.1:8080/embed \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": "Overview of climate change impacts on coastal cities",
    "prompt_name": "document",
  }'

After installing llama.cpp one can run llama-server to host the embedding model as

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys jina-embeddings-text-small-text-matching for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (jina-embeddings-text-small-text-matching below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"jina-embeddings-text-small-text-matching","input":"text to embed"}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms