From the model card
Fields below are copied from the tags and counters on the HuggingFace repository intfloat/multilingual-e5-small at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- intfloat
- Pipeline tag
- sentence-similarity
- Library
- Sentence Transformers
- Framework tags
- PyTorch
- Weight formats
- ONNX, safetensors, OpenVINO
- License tag
mit— read the license file in the repo before relying on it- Language tags
- multilingual; Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Breton (br), Bosnian (bs), Catalan (ca), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Esperanto (eo), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Western Frisian (fy), Irish (ga), Scottish Gaelic (gd), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kurdish (ku), Kyrgyz (ky), Latin (la), Lao (lo), Lithuanian (lt), Latvian (lv), Malagasy (mg), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Oromo (om), Odia (or), Punjabi (pa), Polish (pl), Pashto (ps), Portuguese (pt), Romanian (ro), Russian (ru), Sanskrit (sa), Sindhi (sd), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Uyghur (ug), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Xhosa (xh), Yiddish (yi), Chinese (zh)
- Papers cited
- arXiv:2402.05672, arXiv:2108.08787, arXiv:2104.08663, arXiv:2210.07316
- Downloads (HF counter at last fetch)
- 11,740,779
- Likes (HF counter at last fetch)
- 395
- Model card
- https://huggingface.co/intfloat/multilingual-e5-small
Use cases
- Lightweight multilingual semantic search in resource-constrained environments
- High-throughput multilingual embedding generation at scale
- Cross-lingual retrieval where inference cost matters more than peak accuracy
- Mobile or edge multilingual embedding with CPU inference
- Multilingual RAG embeddings where latency budgets exclude larger models
Pros
- MIT license
- 100+ language coverage in a compact model
- ONNX and OpenVINO compatible; text-embeddings-inference support
- Instruction prefix training for asymmetric retrieval tasks
Cons
- Accuracy below multilingual-e5-large and BGE-M3 on hard multilingual retrieval
- Low-resource language quality gap more pronounced at smaller scale
- Instruction prefix required for best performance
- BERT backbone limits capacity for complex multilingual semantic distinctions
- Superseded by newer multilingual models on MTEB leaderboard
Tags
sentence-transformerspytorchonnxsafetensorsopenvinobertmtebSentence Transformerssentence-similaritymultilingualafamarasazbebgbnbrbs