AI Tools.

Search

sentence similarity by BAAI

bge-m3

BAAI's BGE-M3 embedding model supporting over 100 languages with a unified architecture capable of dense, sparse (lexical), and late-interaction (ColBERT-style) retrieval modes from a single checkpoint. Built on XLM-RoBERTa with large-scale multilingual training, it targets multi-lingual and cross-lingual retrieval where a single model must handle diverse language inputs.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository BAAI/bge-m3 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
BAAI
Pipeline tag
sentence-similarity
Library
Sentence Transformers
Framework tags
PyTorch
Weight formats
ONNX
License tag
mit — read the license file in the repo before relying on it
Papers cited
arXiv:2402.03216, arXiv:2004.04906, arXiv:2106.14807, arXiv:2107.05720, arXiv:2004.12832
Downloads (HF counter at last fetch)
36,725,443
Likes (HF counter at last fetch)
3,466
Model card
https://huggingface.co/BAAI/bge-m3

Use cases

  • Multilingual semantic search across 100+ language corpora
  • Cross-lingual retrieval for international knowledge bases and documentation
  • Hybrid dense+sparse retrieval combining semantic and keyword matching signals
  • Dense passage retrieval in RAG pipelines serving non-English content
  • Large-scale multilingual document indexing

Pros

  • 100+ language coverage eliminates per-language model management overhead
  • Unified dense/sparse/ColBERT outputs enable flexible retrieval strategies
  • MIT license; strong MTEB multilingual leaderboard performance
  • XLM-RoBERTa backbone brings established multilingual pretraining quality

Cons

  • Larger than smaller BGE variants, increasing deployment memory requirements
  • Dense + sparse + ColBERT inference modes add compute overhead over single-mode bi-encoders
  • Quality gaps between high-resource and low-resource language coverage
  • Complex deployment compared to standard single-mode embedding models
  • ONNX export may not cover all retrieval modes

Tags

sentence-transformerspytorchonnxxlm-robertafeature-extractionsentence-similarityarxiv:2402.03216arxiv:2004.04906arxiv:2106.14807arxiv:2107.05720arxiv:2004.12832license:miteval-resultstext-embeddings-inferenceendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure