From the model card
Fields below are copied from the tags and counters on the HuggingFace repository google-bert/bert-base-multilingual-uncased at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- google-bert
- Pipeline tag
- fill-mask
- Library
- Transformers
- Framework tags
- PyTorch, TensorFlow, JAX
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual; Afrikaans (af), Albanian (sq), Arabic (ar), Aragonese (an), Armenian (hy), Asturian (ast), Azerbaijani (az), Bashkir (ba), Basque (eu), Bavarian (bar), Belarusian (be), Bangla (bn), Bosnian (bs), Breton (br), Bulgarian (bg), Burmese (my), Catalan (ca), Cebuano (ceb), Chechen (ce), Chinese (zh), Chuvash (cv), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), Galician (gl), Georgian (ka), German (de), Greek (el), Gujarati (gu), Haitian Creole (ht), Hebrew (he), Hindi (hi), Hungarian (hu), Icelandic (is), Ido (io), Indonesian (id), Irish (ga), Italian (it), Japanese (ja), Javanese (jv), Kannada (kn), Kazakh (kk), Kyrgyz (ky), Korean (ko), Latin (la), Latvian (lv), Lithuanian (lt), Low German (nds), Macedonian (mk), Malagasy (mg), Malay (ms), Malayalam (ml), Marathi (mr), Minangkabau (min), Nepali (ne), Newari (new), Norwegian Bokmål (nb), Norwegian Nynorsk (nn), Occitan (oc), Persian (fa), Piedmontese (pms), Polish (pl), Portuguese (pt), Punjabi (pa), Romanian (ro), Russian (ru), Scots (sco), Serbian (sr), Sicilian (scn), Slovak (sk), Slovenian (sl), Azerbaijani (aze), Spanish (es), Sundanese (su), Swahili (sw), Swedish (sv), Filipino (tl), Tajik (tg), Tamil (ta), Tatar (tt), Telugu (te), Turkish (tr), Ukrainian (uk), Uzbek (uz), Vietnamese (vi), Volapük (vo), Waray (war), Welsh (cy), Western Frisian (fry), Western Panjabi (pnb), Yoruba (yo)
- Papers cited
- arXiv:1810.04805
- Datasets declared
- wikipedia
- Downloads (HF counter at last fetch)
- 4,170,938
- Likes (HF counter at last fetch)
- 159
- Model card
- https://huggingface.co/google-bert/bert-base-multilingual-uncased
Use cases
- Cross-lingual text classification with a single model
- Zero-shot transfer to low-resource languages in the 104-language set
- Multilingual masked language model pretraining baseline
- NER and POS tagging in contexts where case carries no meaning
Pros
- Single model spans 104 languages with a shared multilingual vocabulary
- Apache 2.0 license, widely integrated in community NLP pipelines
- Well-understood baseline with extensive published benchmarks
Cons
- Lowercasing removes signals critical for named entity recognition
- Outperformed on most tasks by XLM-RoBERTa-base and above
- Fixed 512-token context limit with no built-in sliding window support
Tags
transformerspytorchtfjaxsafetensorsbertfill-maskmultilingualafsqaranhyastazbaeubarbebn