Model reference · open weights

Octen-Embedding

Octen-Embedding is an open-weight embedding model from Octen, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings Octen 3 variants 115k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What Octen-Embedding is

Octen-Embedding-8B Octen-Embedding-8B is a text embedding model developed by Octen for semantic search and retrieval tasks. This model is fine-tuned from Qwen/Qwen3-Embedding-8B and supports multiple languages, providing high-quality embeddings for various applications. Key Highlights 🥇 RTEB Leaderboard Champion (as of January 12, 2026) - Octen-Embedding-8B ranks #1 on the RTEB Leaderboard with Mean (Task) score of 0.8045 - Excellent performance on both Public (0.7953) and Private (0.8157) datasets - Demonstrates true generalization capability without overfitting to public benchmarks Industry-Oriented Vertical Domain Expertise - Legal: Legal document retrieval - Finance: Financial reports, Q&A, and personal finance content - Healthcare: Medical Q&A, clinical dialogues, and health consultations - Code: Programming problems, code search, and SQL queries Ultra-Long Context Support - Supports up to 32,768 tokens context length - Suitable for processing long documents in legal, healthcare, and other domains - High-dimensional embedding space for rich semantic representation Multilingual Capability - Supports 100+ languages - Includes various programming languages - Strong multilingual, cross-lingual, and code retrieval capabilities Open Source Model List Model Family Design: - Octen-Embedding-8B: Best performance, RTEB #1, for high-precision retrieval - Octen-Embedding-4B: Best in 4B category, balanced performance and efficiency - Octen-Embedding-0.6B: Lightweight deployment, suitable for edge devices and resource-constrained environments For API access, deployment solutions, and technical documentation, visit octen.ai. Experimental Results RTEB Leaderboard (Overall Performance) Model Details - Base Model: Qwen/Qwen3-Embedding-8B - Model Size: 8B parameters - Max Sequence Length: 40,960 tokens - Embedding Dimension: 4096 - Languages: English, Chinese, and multilingual support - Training Method: LoRA fine-tuning Usage Using Sentence Transformers Using Transformers Recommended Use Cases - Semantic search and information retrieval - Document similarity and clustering - Question answering - Cross-lingual retrieval - Text classification with embeddings Known Issues When e

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerOcten
TypeEmbedding models
Parameters (lead)7.6B
Context40k tokens
Variants3
Runs withsentence-transformers
Based onQwen/Qwen3-Embedding-8B
Released2025-12-23
Popularity115k downloads / month
Likes184
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Octen-Embedding-8B7.6BBF16~17.4 GBWeights ↗
Octen-Embedding-4B4.0BBF16~9.3 GBWeights ↗
Octen-Embedding-0.6B596MBF16~1.4 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys octen-embedding for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (octen-embedding below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"octen-embedding","input":"text to embed"}'

Details

Languages, data & research

Languages

en zh multilingual

Tags

sentence-transformers safetensors qwen3 sentence-similarity feature-extraction embedding text-embedding retrieval en zh multilingual text-embeddings-inference endpoints_compatible deploy:azure

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want Octen-Embedding on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms