AI Tools.

Search

sentence similarity by intfloat

multilingual-e5-small

Multilingual-E5-Small is a compact multilingual embedding model from Microsoft Research supporting 100+ languages on a BERT-based backbone, smaller and faster than the E5-large variant. It uses the same instruction-prefix training approach as E5-large ('query:'/'passage:') for asymmetric retrieval. MIT licensed with ONNX and OpenVINO export.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository intfloat/multilingual-e5-small at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
intfloat
Pipeline tag
sentence-similarity
Library
Sentence Transformers
Framework tags
PyTorch
Weight formats
ONNX, safetensors, OpenVINO
License tag
mit — read the license file in the repo before relying on it
Language tags
multilingual; Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Breton (br), Bosnian (bs), Catalan (ca), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Esperanto (eo), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Western Frisian (fy), Irish (ga), Scottish Gaelic (gd), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kurdish (ku), Kyrgyz (ky), Latin (la), Lao (lo), Lithuanian (lt), Latvian (lv), Malagasy (mg), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Oromo (om), Odia (or), Punjabi (pa), Polish (pl), Pashto (ps), Portuguese (pt), Romanian (ro), Russian (ru), Sanskrit (sa), Sindhi (sd), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Uyghur (ug), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Xhosa (xh), Yiddish (yi), Chinese (zh)
Papers cited
arXiv:2402.05672, arXiv:2108.08787, arXiv:2104.08663, arXiv:2210.07316
Downloads (HF counter at last fetch)
11,740,779
Likes (HF counter at last fetch)
395
Model card
https://huggingface.co/intfloat/multilingual-e5-small

Use cases

  • Lightweight multilingual semantic search in resource-constrained environments
  • High-throughput multilingual embedding generation at scale
  • Cross-lingual retrieval where inference cost matters more than peak accuracy
  • Mobile or edge multilingual embedding with CPU inference
  • Multilingual RAG embeddings where latency budgets exclude larger models

Pros

  • MIT license
  • 100+ language coverage in a compact model
  • ONNX and OpenVINO compatible; text-embeddings-inference support
  • Instruction prefix training for asymmetric retrieval tasks

Cons

  • Accuracy below multilingual-e5-large and BGE-M3 on hard multilingual retrieval
  • Low-resource language quality gap more pronounced at smaller scale
  • Instruction prefix required for best performance
  • BERT backbone limits capacity for complex multilingual semantic distinctions
  • Superseded by newer multilingual models on MTEB leaderboard

Tags

sentence-transformerspytorchonnxsafetensorsopenvinobertmtebSentence Transformerssentence-similaritymultilingualafamarasazbebgbnbrbs