From the model card
Fields below are copied from the tags and counters on the HuggingFace repository jinaai/jina-embeddings-v3 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- jinaai
- Pipeline tag
- feature-extraction
- Library
- Transformers, Sentence Transformers
- Framework tags
- PyTorch
- Weight formats
- ONNX, safetensors
- License tag
cc-by-nc-4.0— read the license file in the repo before relying on it- Language tags
- multilingual; Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bulgarian (bg), Bangla (bn), Breton (br), Bosnian (bs), Catalan (ca), Czech (cs), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Esperanto (eo), Spanish (es), Estonian (et), Basque (eu), Persian (fa), Finnish (fi), French (fr), Western Frisian (fy), Irish (ga), Scottish Gaelic (gd), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Indonesian (id), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Khmer (km), Kannada (kn), Korean (ko), Kurdish (ku), Kyrgyz (ky), Latin (la), Lao (lo), Lithuanian (lt), Latvian (lv), Malagasy (mg), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch (nl), Norwegian (no), Oromo (om), Odia (or), Punjabi (pa), Polish (pl), Pashto (ps), Portuguese (pt), Romanian (ro), Russian (ru), Sanskrit (sa), Sindhi (sd), Sinhala (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Thai (th), Filipino (tl), Turkish (tr), Uyghur (ug), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Xhosa (xh), Yiddish (yi), Chinese (zh)
- Papers cited
- arXiv:2409.10173
- Downloads (HF counter at last fetch)
- 2,319,256
- Likes (HF counter at last fetch)
- 1,154
- Model card
- https://huggingface.co/jinaai/jina-embeddings-v3
Use cases
- Long-document retrieval where 512-token context is insufficient
- Multilingual semantic search across 89 languages
- Task-adaptive embeddings that switch modes based on query type
- Replacing multiple task-specific embedding models with one deployment
Pros
- 8192-token context window — far beyond most open embedding models
- Single model covers retrieval, classification, and similarity via LoRA adapters
- Strong MTEB multilingual scores
- CC-BY-NC 4.0 — free for non-commercial use
Cons
- CC-BY-NC license prohibits commercial use without a Jina AI license
- 570M parameters require ~1.2GB VRAM — heavier than MiniLM-class models
- LoRA adapter switching adds configuration complexity
- Commercial users must contact Jina AI for licensing terms
Tags
transformerspytorchonnxsafetensorsfeature-extractionsentence-similaritymtebsentence-transformerscustom_codemultilingualafamarasazbebgbnbrbs