AI Tools.

Search

feature extraction by BAAI

bge-large-en-v1.5

BGE-Large-EN-v1.5 is BAAI's highest-capacity English embedding model in the v1.5 series, producing 1024-dimensional vectors. It achieves top MTEB retrieval scores among its generation of English-only embedding models, at the cost of higher compute and storage than BGE-small or BGE-base. MIT licensed with ONNX export support.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository BAAI/bge-large-en-v1.5 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
BAAI
Pipeline tag
feature-extraction
Library
Sentence Transformers, Transformers
Framework tags
PyTorch
Weight formats
ONNX, safetensors
License tag
mit — read the license file in the repo before relying on it
Language tags
English (en)
Papers cited
arXiv:2401.03462, arXiv:2312.15503, arXiv:2311.13534, arXiv:2310.07554, arXiv:2309.07597
Downloads (HF counter at last fetch)
13,118,409
Likes (HF counter at last fetch)
720
Model card
https://huggingface.co/BAAI/bge-large-en-v1.5

Use cases

  • High-precision semantic search where embedding quality is the primary constraint
  • Embedding for legal, medical, or technical domain retrieval requiring fine-grained distinction
  • MTEB benchmark baseline as a strong English embedding reference point
  • Re-ranking large candidate sets using embedding similarity
  • Knowledge base retrieval where 768-dim models underperform

Pros

  • Strong MTEB retrieval accuracy at 1024 dimensions
  • MIT license for commercial use
  • ONNX and text-embeddings-inference compatible for production deployment
  • Part of the well-maintained BAAI BGE family with documented benchmarks

Cons

  • 1024-dim output doubles storage cost vs. 512-dim alternatives
  • Higher inference compute than BGE-small or BGE-base
  • English-only; no multilingual or cross-lingual capability
  • May provide marginal gains over BGE-base for many standard retrieval tasks
  • Newer instruction-following embedding models are competitive at smaller sizes

Tags

sentence-transformerspytorchonnxsafetensorsbertfeature-extractionsentence-similaritytransformersmtebenarxiv:2401.03462arxiv:2312.15503arxiv:2311.13534arxiv:2310.07554arxiv:2309.07597license:mitmodel-indexeval-resultstext-embeddings-inferenceendpoints_compatible