From the model card
Fields below are copied from the tags and counters on the HuggingFace repository BAAI/bge-large-en-v1.5 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- BAAI
- Pipeline tag
- feature-extraction
- Library
- Sentence Transformers, Transformers
- Framework tags
- PyTorch
- Weight formats
- ONNX, safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:2401.03462, arXiv:2312.15503, arXiv:2311.13534, arXiv:2310.07554, arXiv:2309.07597
- Downloads (HF counter at last fetch)
- 13,118,409
- Likes (HF counter at last fetch)
- 720
- Model card
- https://huggingface.co/BAAI/bge-large-en-v1.5
Use cases
- High-precision semantic search where embedding quality is the primary constraint
- Embedding for legal, medical, or technical domain retrieval requiring fine-grained distinction
- MTEB benchmark baseline as a strong English embedding reference point
- Re-ranking large candidate sets using embedding similarity
- Knowledge base retrieval where 768-dim models underperform
Pros
- Strong MTEB retrieval accuracy at 1024 dimensions
- MIT license for commercial use
- ONNX and text-embeddings-inference compatible for production deployment
- Part of the well-maintained BAAI BGE family with documented benchmarks
Cons
- 1024-dim output doubles storage cost vs. 512-dim alternatives
- Higher inference compute than BGE-small or BGE-base
- English-only; no multilingual or cross-lingual capability
- May provide marginal gains over BGE-base for many standard retrieval tasks
- Newer instruction-following embedding models are competitive at smaller sizes
Tags
sentence-transformerspytorchonnxsafetensorsbertfeature-extractionsentence-similaritytransformersmtebenarxiv:2401.03462arxiv:2312.15503arxiv:2311.13534arxiv:2310.07554arxiv:2309.07597license:mitmodel-indexeval-resultstext-embeddings-inferenceendpoints_compatible