Model reference · open weights

bge-large-en

bge-large-en is an open-weight embedding model from BAAI, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings BAAI 1 variants 13.8M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What bge-large-en is

For more details please refer to our Github: FlagEmbedding. If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3. English | 中文 FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: - Long-Context LLM: Activation Beacon - Fine-tuning of LM : LM-Cocktail - Dense Retrieval: BGE-M3, LLM Embedder, BGE Embedding - Reranker Model: BGE Reranker - Benchmark: C-MTEB News - 1/30/2024: Release BGE-M3, a new member to BGE model series! M3 stands for Multi-linguality (100+ languages), Multi-granularities (input length up to 8192), Multi-Functionality (unification of dense, lexical, multi-vec/colbert retrieval). It is the first embedding model that supports all three retrieval methods, achieving new SOTA on multi-lingual (MIRACL) and cross-lingual (MKQA) benchmarks. Technical Report and Code. :fire: - 1/9/2024: Release Activation-Beacon, an effective, efficient, compatible, and low-cost (training) method to extend the context length of LLM. Technical Report :fire: - 12/24/2023: Release LLaRA, a LLaMA-7B based dense retriever, leading to state-of-the-art performances on MS MARCO and BEIR. Model and code will be open-sourced. Please stay tuned. Technical Report :fire: - 11/23/2023: Release LM-Cocktail, a method to maintain general capabilities during fine-tuning by merging multiple language models. Technical Report :fire: - 10/12/2023: Release LLM-Embedder, a unified embedding model to support diverse retrieval augmentation needs for LLMs. Technical Report - 09/15/2023: The technical report and massive training data of BGE has been released - 09/12/2023: New models: - New reranker model: release cross-encoder models BAAI/bge-reranker-base and BAAI/bge-reranker-large, which are more powerful than embedding model. We recommend to use/fine-tune them to re-rank top-k documents returned by embedding models. - update embedding model: release bge--v1.5 embedding model to alleviate the issue of the similarity distribution, and enhance its retrieval ability without instruction. - 09/07/2023: Update fine-tune code: Add script to mine hard negatives and support adding in

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerBAAI
TypeEmbedding models
Parameters (lead)335M
Context512 tokens
Variants1
Runs withsentence-transformers
Released2023-09-12
Popularity13.8M downloads / month
Likes718
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
bge-large-en-v1.5335MBF16~0.8 GBWeights ↗

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
ClassificationMTEB AmazonCounterfactualClassification (en)accuracy75.851
ClassificationMTEB AmazonCounterfactualClassification (en)ap38.566
ClassificationMTEB AmazonCounterfactualClassification (en)f169.694
ClassificationMTEB AmazonPolarityClassificationaccuracy92.417
ClassificationMTEB AmazonPolarityClassificationap89.193
ClassificationMTEB AmazonPolarityClassificationf192.395
ClassificationMTEB AmazonReviewsClassification (en)accuracy48.176
ClassificationMTEB AmazonReviewsClassification (en)f147.807
RetrievalMTEB ArguAnamap_at_140.185
RetrievalMTEB ArguAnamap_at_1055.654
RetrievalMTEB ArguAnamap_at_10056.25
RetrievalMTEB ArguAnamap_at_100056.255
RetrievalMTEB ArguAnamap_at_351.743
RetrievalMTEB ArguAnamap_at_554.129
RetrievalMTEB ArguAnamrr_at_140.967
RetrievalMTEB ArguAnamrr_at_1055.96
RetrievalMTEB ArguAnamrr_at_10056.549
RetrievalMTEB ArguAnamrr_at_100056.554
RetrievalMTEB ArguAnamrr_at_351.98
RetrievalMTEB ArguAnamrr_at_554.44
RetrievalMTEB ArguAnandcg_at_140.185
RetrievalMTEB ArguAnandcg_at_1063.542
RetrievalMTEB ArguAnandcg_at_10065.965
RetrievalMTEB ArguAnandcg_at_100066.087

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys bge-large-en for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (bge-large-en below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-large-en","input":"text to embed"}'

Details

Languages, data & research

Languages

en

Tags

sentence-transformers pytorch onnx safetensors bert feature-extraction sentence-similarity transformers mteb en model-index eval-results text-embeddings-inference endpoints_compatible

Papers

Licence

Open weights

Open weights under mit — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want bge-large-en on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms