Model reference · open weights

vietnamese-bi-encoder

vietnamese-bi-encoder is an open-weight embedding model from bkai-foundation-models, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings bkai-foundation-models 1 variants 459k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What vietnamese-bi-encoder is

bkai-foundation-models/vietnamese-bi-encoder This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. We train the model on a merged training dataset that consists of: - MS Macro (translated into Vietnamese) - SQuAD v2 (translated into Vietnamese) - 80% of the training set from the Legal Text Retrieval Zalo 2021 challenge We use phobert-base-v2 as the pre-trained backbone. Here are the results on the remaining 20% of the training set from the Legal Text Retrieval Zalo 2021 challenge: Usage (Sentence-Transformers) Using this model becomes easy when you have sentence-transformers installed: Then you can use the model like this: Usage (Widget HuggingFace) The widget use custom pipeline on top of the default pipeline by adding additional word segmenter before PhobertTokenizer. So you do not need to segment words before using the API: An example could be seen in Hosted inference API. Usage (HuggingFace Transformers) Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings. Training The model was trained with the parameters: DataLoader: torch.utils.data.dataloader.DataLoader of length 17584 with parameters: Loss: sentencetransformers.losses.MultipleNegativesRankingLoss.MultipleNegativesRankingLoss with parameters: Parameters of the fit()-Method: Full Model Architecture Please cite our manuscript if this dataset is used for your work

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makerbkai-foundation-models
TypeEmbedding models
Parameters (lead)135M
Context258 tokens
Variants1
Runs withgeneric
Released2023-09-09
Popularity459k downloads / month
Likes79
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
vietnamese-bi-encoder135MBF16~0.3 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys vietnamese-bi-encoder for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (vietnamese-bi-encoder below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"vietnamese-bi-encoder","input":"text to embed"}'

Details

Languages, data & research

Languages

vi

Tags

generic pytorch safetensors roberta feature-extraction sentence-transformers sentence-similarity transformers vi

Papers

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want vietnamese-bi-encoder on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms