From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Alibaba-NLP/gte-multilingual-base at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Alibaba-NLP
- Pipeline tag
- sentence-similarity
- Library
- Sentence Transformers, Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual; Newari (new), Afrikaans (af), Arabic (ar), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Catalan (ca), Cebuano (ceb), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Galician (gl), Gujarati (gu), Hebrew (he), Hindi (hi), Croatian (hr), Haitian Creole (ht), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kyrgyz (ky), Lao (lo), Lithuanian (lt), Latvian (lv), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Punjabi (pa), Polish (pl), Portuguese (pt), Quechua (qu), Romanian (ro), Russian (ru), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Ukrainian (uk), Urdu (ur), Vietnamese (vi), Yoruba (yo), Chinese (zh)
- Papers cited
- arXiv:2407.19669, arXiv:2210.09984, arXiv:2402.03216, arXiv:2007.15207, arXiv:2104.08663, arXiv:2402.07440
- Downloads (HF counter at last fetch)
- 1,297,884
- Likes (HF counter at last fetch)
- 375
- Model card
- https://huggingface.co/Alibaba-NLP/gte-multilingual-base
Use cases
- Multilingual semantic search across mixed-language corpora
- Cross-lingual document retrieval without separate models per language
- Building multilingual RAG pipelines for global applications
- Multilingual sentence similarity benchmarking
Pros
- Strong MTEB multilingual scores competitive with E5-multilingual-large
- Apache-2.0 licensed
- Single model covers 70+ languages without per-language tuning
- 768-dim output compatible with existing vector database deployments
Cons
- 305M parameters are significantly heavier than MiniLM-class models
- Lower quality than language-specific models on high-resource languages like Chinese and German
- Limited documentation on which languages were in training data
- Requires careful normalization setup for best retrieval results
Tags
sentence-transformerssafetensorsnewfeature-extractionmtebtransformersmultilingualsentence-similaritytext-embeddings-inferencecustom_codeafarazbebgbncacebcscy