Model reference · open weights
bekko-embedding is an open-weight embedding model from hotchpotch. bekko-embedding-v1-a8m (BF16) weighs 212 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | hotchpotch |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 123M |
| Context | 8,192 tokens |
| Runs with | sentence-transformers |
| Based on | hotchpotch/bekko-embedding-v1-a25m-pt |
| Released | 2026-07-19 |
| Popularity | 16k downloads / month |
| Weights | 212 MB (bekko-embedding-v1-a8m (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 212 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
Builds
| Build | Parameters | Precision | Weights | Smallest card |
|---|---|---|---|---|
| bekko-embedding-v1-a25m ↗ | 123M | FP32 | 246 MB | RTX 3060 12 GB |
| bekko-embedding-v1-a8m (above) ↗ | 106M | BF16 | 212 MB | RTX 3060 12 GB |
Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.
From the model card
bekko-embedding-v1-a25m is an ultra-compact multilingual text embedding model. It has just 25M active parameters — light enough to run comfortably on modest CPUs — yet its retrieval quality is comparable to models with 3–10x more active parameters.
For a smaller, faster model, see bekko-embedding-v1-a8m (8M active parameters).
You can also try bekko right in your browser: the bekko-embedding-web demo runs the model fully client-side with Transformers.js — no server involved.
[!NOTE] For a guided overview of the models, training recipe, and results, read Bekko Embedding: how small can a multilingual retrieval model be?.
| a8m | a25m (this model) | |
|---|---|---|
| Active parameters | 7.7M | 24.9M |
| HAKARI-Bench overall | 0.545 | 0.570 |
| MMTEB Retrieval | 56.2 | 57.5 |
| CPU docs/s (Ryzen 9 7950X, OpenVINO) | 364 | 134 |
| CPU docs/s (Raspberry Pi 5, OpenVINO) | 33 | 10.5 |
| GPU docs/s (RTX 5090, Flash Attention 2) | 5,561 | 4,006 |
Rule of thumb: a25m is the quality pick. Switch to a8m when CPU budget or latency is tight — it keeps most of the quality and gains about 2.7x CPU throughput.
We recommend Sentence Transformers 5.0+ and Transformers 5.12+:
pip install -U "sentence-transformers>=5.0" "transformers>=5.12"
Queries and documents go through the same encode() call — no prefixes or task instructions needed. Pass normalize_embeddings=True when you plan to search with cosine similarity or dot product.
On GPU, SDPA works out of the box with PyTorch and CUDA. Flash Attention 2 requires pip install flash-attn --no-build-isolation; on our RTX 5090 it was about 24% faster, and can be enabled by replacing "sdpa" below with "flash_attention_2". Sentence Transformers selects CUDA automatically, so device is normally unnecessary; to force it, use device="cuda", not "gpu".
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer(
"hotchpotch/bekko-embedding-v1-a25m",
# model_kwargs={"attn_implementation": "sdpa"}, # Optional on GPU
)
query = "What are the characteristics of sushi?"
docs = [
"A warm noodle soup served in broth with sliced toppings.",
"天ぷらは魚や野菜に衣をつけて揚げた料理です。", # "Tempura is battered, deep-fried fish and vegetables."
"Une fine crepe garnie de sucre, de beurre ou de fruits.",
"A Japanese dish made with vinegared rice, often shaped with seafood, vegetables, or egg.",
]
query_emb = model.encode(query, normalize_embeddings=True)
doc_emb = model.encode(docs, normalize_embeddings=True)
scores = util.cos_sim(query_emb, doc_emb)[0]
print(scores)
print("best doc:", docs[int(scores.argmax())])
Output (exact scores vary slightly by backend):
tensor([0.2953, 0.2785, 0.3209, 0.4378])
best doc: A Japanese dish made with vinegared rice, often shaped with seafood, vegetables, or egg.
Queries and documents don't need to share a language. Continuing with the same model, a Japanese query finds the right English document in a mixed English / Spanish corpus:
corpus = [
"Sushi is a Japanese dish of vinegared rice topped with seafood.",
"The Eiffel Tower is a wrought-iron lattice tower in Paris, France.",
"Mount Fuji is the highest mountain in Japan, at 3,776 meters.",
"Python is a programming language known for its readability.",
"La Sagrada Família es una basílica de Barcelona diseñada por Antoni Gaudí.",
]
corpus_emb = model.encode(corpus, normalize_embeddings=True)
for query in [
"日本で一番高い山は?", # "What is the highest mountain in Japan?"
"Who designed the famous basilica in Barcelona?",
]:
query_emb = model.encode(query, normalize_embeddings=True)
hits = util.semantic_search(query_emb, corpus_emb, top_k=2)[0]
print(query)
for hit in hits:
print(f" {hit['score']:.3f} {corpus[hit['corpus_id']]}")
日本で一番高い山は?
0.457 Mount Fuji is the highest mountain in Japan, at 3,776 meters.
0.142 Sushi is a Japanese dish of vinegared rice topped with seafood.
Who designed the famous basilica in Barcelona?
0.563 La Sagrada Família es una basílica de Barcelona diseñada por Antoni Gaudí.
0.126 The Eiffel Tower is a wrought-iron lattice tower in Paris, France.
That's everything you need for basic use. For more speed — OpenVINO on CPU, Flash Attention on GPU, browser inference, smaller embeddings — see Optimized Inference below.
In the chart above, up and to the left is better: more retrieval quality from fewer active parameters. The step line shows the best observed score within each active-parameter budget, and outlined markers identify Pareto-efficient models. Both bekko models sit in that upper-left region, scoring at or above many models several times their size — which is the whole point of the project.
On the 131-task MMTEB Multilingual v2 suite, a25m scores 57.5 Retrieval and 58.3 Mean(Task) with 24.9M active parameters. That edges out gte-multilingual-base on Retrieval and ties it on Mean with ~4.5x fewer active parameters, and beats multilingual-e5-large and BGE-M3 on Retrieval with ~12x fewer.
Scores are ×100. Retrieval is task-macro nDCG@10, and Mean is the mean across all 131 tasks. Competitor values use the official 2026-06-28 snapshot.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.