Model reference · open weights
vietnamese-bi-encoder is an open-weight embedding model from bkai-foundation-models, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.
About
bkai-foundation-models/vietnamese-bi-encoder This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. We train the model on a merged training dataset that consists of: - MS Macro (translated into Vietnamese) - SQuAD v2 (translated into Vietnamese) - 80% of the training set from the Legal Text Retrieval Zalo 2021 challenge We use phobert-base-v2 as the pre-trained backbone. Here are the results on the remaining 20% of the training set from the Legal Text Retrieval Zalo 2021 challenge: Usage (Sentence-Transformers) Using this model becomes easy when you have sentence-transformers installed: Then you can use the model like this: Usage (Widget HuggingFace) The widget use custom pipeline on top of the default pipeline by adding additional word segmenter before PhobertTokenizer. So you do not need to segment words before using the API: An example could be seen in Hosted inference API. Usage (HuggingFace Transformers) Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings. Training The model was trained with the parameters: DataLoader: torch.utils.data.dataloader.DataLoader of length 17584 with parameters: Loss: sentencetransformers.losses.MultipleNegativesRankingLoss.MultipleNegativesRankingLoss with parameters: Parameters of the fit()-Method: Full Model Architecture Please cite our manuscript if this dataset is used for your work
Summarised from the published model card. Read the full card on the HuggingFace links below.
Specifications
| Maker | bkai-foundation-models |
|---|---|
| Type | Embedding models |
| Parameters (lead) | 135M |
| Context | 258 tokens |
| Variants | 1 |
| Runs with | generic |
| Released | 2023-09-09 |
| Popularity | 459k downloads / month |
| Likes | 79 |
| Licence | Open weights |
How it works
Variants
Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.
| Variant | Params | Precision | VRAM | Fits 16 GB | Weights |
|---|---|---|---|---|---|
| vietnamese-bi-encoder | 135M | BF16 | ~0.3 GB | ✓ | Weights ↗ |
Using it via the API
Once AxForge deploys vietnamese-bi-encoder for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (vietnamese-bi-encoder below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"vietnamese-bi-encoder","input":"text to embed"}'
Licence
Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗
Explore