AI Tools.

Search

sentence similarity by Alibaba-NLP

gte-multilingual-base

GTE-multilingual-base is Alibaba's 305M-parameter embedding model covering 70+ languages, designed for multilingual dense retrieval and semantic similarity. It uses a modified transformer backbone with improved positional encoding for cross-lingual transfer.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Alibaba-NLP/gte-multilingual-base at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Alibaba-NLP
Pipeline tag
sentence-similarity
Library
Sentence Transformers, Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual; Newari (new), Afrikaans (af), Arabic (ar), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Catalan (ca), Cebuano (ceb), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Galician (gl), Gujarati (gu), Hebrew (he), Hindi (hi), Croatian (hr), Haitian Creole (ht), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kyrgyz (ky), Lao (lo), Lithuanian (lt), Latvian (lv), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Punjabi (pa), Polish (pl), Portuguese (pt), Quechua (qu), Romanian (ro), Russian (ru), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Ukrainian (uk), Urdu (ur), Vietnamese (vi), Yoruba (yo), Chinese (zh)
Papers cited
arXiv:2407.19669, arXiv:2210.09984, arXiv:2402.03216, arXiv:2007.15207, arXiv:2104.08663, arXiv:2402.07440
Downloads (HF counter at last fetch)
1,297,884
Likes (HF counter at last fetch)
375
Model card
https://huggingface.co/Alibaba-NLP/gte-multilingual-base

Use cases

  • Multilingual semantic search across mixed-language corpora
  • Cross-lingual document retrieval without separate models per language
  • Building multilingual RAG pipelines for global applications
  • Multilingual sentence similarity benchmarking

Pros

  • Strong MTEB multilingual scores competitive with E5-multilingual-large
  • Apache-2.0 licensed
  • Single model covers 70+ languages without per-language tuning
  • 768-dim output compatible with existing vector database deployments

Cons

  • 305M parameters are significantly heavier than MiniLM-class models
  • Lower quality than language-specific models on high-resource languages like Chinese and German
  • Limited documentation on which languages were in training data
  • Requires careful normalization setup for best retrieval results

Tags

sentence-transformerssafetensorsnewfeature-extractionmtebtransformersmultilingualsentence-similaritytext-embeddings-inferencecustom_codeafarazbebgbncacebcscy