From the model card
Fields below are copied from the tags and counters on the HuggingFace repository tohoku-nlp/bert-base-japanese-whole-word-masking at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- tohoku-nlp
- Pipeline tag
- fill-mask
- Library
- Transformers
- Framework tags
- PyTorch, TensorFlow, JAX
- License tag
cc-by-sa-4.0— read the license file in the repo before relying on it- Language tags
- Japanese (ja)
- Datasets declared
- wikipedia
- Downloads (HF counter at last fetch)
- 357,704
- Likes (HF counter at last fetch)
- 76
- Model card
- https://huggingface.co/tohoku-nlp/bert-base-japanese-whole-word-masking
Use cases
- Japanese text classification (sentiment, category, intent)
- Named entity recognition in Japanese documents
- Japanese semantic similarity and sentence embedding with fine-tuning
- Foundation model for Japanese NLP fine-tuning experiments
Pros
- Whole-word masking aligns better with Japanese morphology than character masking
- Widely used in Japanese NLP research — comparable results available in literature
- Maintained by Tohoku NLP, an active Japanese NLP group
- CC BY-SA 4.0 license
Cons
- Japanese-only — not useful for multilingual tasks
- Outperformed by larger models (DeBERTa-v3-base-japanese, multilingual alternatives) on modern benchmarks
- 512-token BERT context limit may truncate longer Japanese documents
- Wikipedia-only pretraining biases toward encyclopedic formal text
Tags
transformerspytorchtfjaxbertfill-maskjadataset:wikipedialicense:cc-by-sa-4.0endpoints_compatibleregion:usdeploy:azure