AI Tools.

Search

fill mask by google-bert

bert-base-multilingual-uncased

BERT-base-multilingual-uncased is Google's multilingual BERT trained on Wikipedia text from 104 languages with all text lowercased before tokenization. Lowercasing simplifies processing but removes capitalization signals that help named entity recognition. It produces 768-dimensional embeddings shared across all supported languages.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository google-bert/bert-base-multilingual-uncased at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
google-bert
Pipeline tag
fill-mask
Library
Transformers
Framework tags
PyTorch, TensorFlow, JAX
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual; Afrikaans (af), Albanian (sq), Arabic (ar), Aragonese (an), Armenian (hy), Asturian (ast), Azerbaijani (az), Bashkir (ba), Basque (eu), Bavarian (bar), Belarusian (be), Bangla (bn), Bosnian (bs), Breton (br), Bulgarian (bg), Burmese (my), Catalan (ca), Cebuano (ceb), Chechen (ce), Chinese (zh), Chuvash (cv), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), Galician (gl), Georgian (ka), German (de), Greek (el), Gujarati (gu), Haitian Creole (ht), Hebrew (he), Hindi (hi), Hungarian (hu), Icelandic (is), Ido (io), Indonesian (id), Irish (ga), Italian (it), Japanese (ja), Javanese (jv), Kannada (kn), Kazakh (kk), Kyrgyz (ky), Korean (ko), Latin (la), Latvian (lv), Lithuanian (lt), Low German (nds), Macedonian (mk), Malagasy (mg), Malay (ms), Malayalam (ml), Marathi (mr), Minangkabau (min), Nepali (ne), Newari (new), Norwegian Bokmål (nb), Norwegian Nynorsk (nn), Occitan (oc), Persian (fa), Piedmontese (pms), Polish (pl), Portuguese (pt), Punjabi (pa), Romanian (ro), Russian (ru), Scots (sco), Serbian (sr), Sicilian (scn), Slovak (sk), Slovenian (sl), Azerbaijani (aze), Spanish (es), Sundanese (su), Swahili (sw), Swedish (sv), Filipino (tl), Tajik (tg), Tamil (ta), Tatar (tt), Telugu (te), Turkish (tr), Ukrainian (uk), Uzbek (uz), Vietnamese (vi), Volapük (vo), Waray (war), Welsh (cy), Western Frisian (fry), Western Panjabi (pnb), Yoruba (yo)
Papers cited
arXiv:1810.04805
Datasets declared
wikipedia
Downloads (HF counter at last fetch)
4,170,938
Likes (HF counter at last fetch)
159
Model card
https://huggingface.co/google-bert/bert-base-multilingual-uncased

Use cases

  • Cross-lingual text classification with a single model
  • Zero-shot transfer to low-resource languages in the 104-language set
  • Multilingual masked language model pretraining baseline
  • NER and POS tagging in contexts where case carries no meaning

Pros

  • Single model spans 104 languages with a shared multilingual vocabulary
  • Apache 2.0 license, widely integrated in community NLP pipelines
  • Well-understood baseline with extensive published benchmarks

Cons

  • Lowercasing removes signals critical for named entity recognition
  • Outperformed on most tasks by XLM-RoBERTa-base and above
  • Fixed 512-token context limit with no built-in sliding window support

Tags

transformerspytorchtfjaxsafetensorsbertfill-maskmultilingualafsqaranhyastazbaeubarbebn