AI Tools.

Search

feature extraction by intfloat

multilingual-e5-large

Multilingual-E5-Large is a 560-million-parameter multilingual embedding model from Microsoft Research, supporting 100+ languages via an XLM-RoBERTa backbone. Trained with E5's instruction-following approach (prepending 'query:' or 'passage:' prefixes), it achieves strong MTEB multilingual retrieval scores. MIT licensed with ONNX and OpenVINO export.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository intfloat/multilingual-e5-large at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
intfloat
Pipeline tag
feature-extraction
Library
Sentence Transformers
Framework tags
PyTorch
Weight formats
ONNX, safetensors, OpenVINO
License tag
mit — read the license file in the repo before relying on it
Language tags
multilingual; Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Breton (br), Bosnian (bs), Catalan (ca), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Esperanto (eo), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Western Frisian (fy), Irish (ga), Scottish Gaelic (gd), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kurdish (ku), Kyrgyz (ky), Latin (la), Lao (lo), Lithuanian (lt), Latvian (lv), Malagasy (mg), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Oromo (om), Odia (or), Punjabi (pa), Polish (pl), Pashto (ps), Portuguese (pt), Romanian (ro), Russian (ru), Sanskrit (sa), Sindhi (sd), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Uyghur (ug), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Xhosa (xh), Yiddish (yi), Chinese (zh)
Papers cited
arXiv:2402.05672, arXiv:2108.08787, arXiv:2104.08663, arXiv:2210.07316
Downloads (HF counter at last fetch)
6,953,891
Likes (HF counter at last fetch)
1,245
Model card
https://huggingface.co/intfloat/multilingual-e5-large

Use cases

  • Multilingual semantic search across 100-language corpora
  • Cross-lingual retrieval where query and documents are in different languages
  • Multilingual RAG pipeline embedding for international content
  • Dense retrieval for low-resource language content with cross-lingual transfer
  • Multilingual text clustering and classification via embeddings

Pros

  • MIT license for commercial use
  • 100+ language coverage with strong multilingual retrieval performance
  • Instruction prefix support ('query:'/'passage:') for asymmetric retrieval
  • ONNX and OpenVINO export; text-embeddings-inference compatible

Cons

  • 560M parameters make it significantly heavier than lighter multilingual models (BGE-M3-small)
  • Larger model size requires more VRAM for batch inference than BGE-M3 or paraphrase-multilingual-MiniLM
  • Quality varies for low-resource languages despite 100+ coverage
  • Instruction prefix is required for best performance — models without the prefix produce degraded embeddings
  • Less adopted than BGE-M3 in the multilingual embedding community

Tags

sentence-transformerspytorchonnxsafetensorsopenvinoxlm-robertamtebSentence Transformerssentence-similarityfeature-extractionmultilingualafamarasazbebgbnbr