by BAAI
BGE-Large-EN-v1.5 is BAAI's highest-capacity English embedding model in the v1.5 series, producing 1024-dimensional vectors. It achieves top MTEB retrieval scores among its generation of English-only embedding models, at the cost of higher compute and storage than BGE-small or BGE-base. MIT licensed with ONNX export support.
13,118,409 ↓ · 720 ♡
by intfloat
Multilingual-E5-Large is a 560-million-parameter multilingual embedding model from Microsoft Research, supporting 100+ languages via an XLM-RoBERTa backbone. Trained with E5's instruction-following approach (prepending 'query:' or 'passage:' prefixes), it achieves strong MTEB multilingual retrieval scores. MIT licensed with ONNX and OpenVINO export.
6,953,891 ↓ · 1,245 ♡
by Qwen
Qwen3-Embedding-0.6B is Alibaba Cloud's compact embedding model from the Qwen3 series, fine-tuned from Qwen3-0.6B-Base for text embedding tasks. At 0.6B parameters it provides instruction-following embedding capability at a size deployable without dedicated GPU infrastructure. Apache 2.0 licensed.
6,612,384 ↓ · 1,179 ♡
by jinaai
Jina Embeddings v3 is a 570M-parameter text embedding model supporting 89 languages with a 8192-token context window. It uses LoRA adapters to switch between task-specific embedding modes (retrieval, similarity, classification) without separate models.
2,319,256 ↓ · 1,154 ♡
by facebook
Meta's wav2vec-BERT 2.0 is a self-supervised speech encoder that combines contrastive learning with masked language modeling objectives. It serves as the backbone for Seamless and other Meta speech recognition and translation systems.
2,245,550 ↓ · 227 ♡
by intfloat
E5-Mistral-7B-Instruct is an embedding model that leverages the full generative capacity of Mistral 7B by using decoder-only LLM representations for text embeddings. It uses instruction prompts at inference time to orient embeddings for retrieval, clustering, or classification tasks. At release it achieved state-of-the-art MTEB scores for dense retrieval, outperforming BERT-family embedding models by a significant margin on hard retrieval tasks.
520,252 ↓ · 569 ♡
by ai-forever
RoSBERTa is a bilingual Russian-English sentence embedding model from ai-forever, built on RoBERTa with MTEB-style training for semantic similarity. It targets retrieval and semantic search use cases in Russian-language NLP pipelines. MIT-licensed and available with text-embeddings-inference compatibility.
511,825 ↓ · 82 ♡
by unsloth
Unsloth's Sentence Transformers conversion of Alibaba's Qwen3-Embedding-4B, enabling direct use for semantic similarity and retrieval. The 4B Qwen3-based embedding model (arXiv:2506.05176) achieves strong MTEB retrieval scores at the cost of a larger memory footprint than typical embedding models.
343,270 ↓ · 2 ♡