From the model card
Fields below are copied from the tags and counters on the HuggingFace repository neuralmind/bert-large-portuguese-cased at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- neuralmind
- Pipeline tag
- fill-mask
- Library
- Transformers
- Framework tags
- PyTorch, JAX
- License tag
mit— read the license file in the repo before relying on it- Language tags
- Portuguese (pt)
- Datasets declared
- brWaC
- Downloads (HF counter at last fetch)
- 1,561,757
- Likes (HF counter at last fetch)
- 74
- Model card
- https://huggingface.co/neuralmind/bert-large-portuguese-cased
Use cases
- Portuguese text classification and entity recognition
- Sentiment analysis on Portuguese social media and news
- Portuguese question answering and reading comprehension
- Fine-tuning base for Brazilian Portuguese NLP applications
Pros
- Pretrained on Portuguese text — far outperforms multilingual BERT on Portuguese tasks
- Apache-2.0 licensed
- Cased variant preserves proper noun information
- Well-benchmarked on standard Portuguese NLP datasets
Cons
- Portuguese-only — not useful for bilingual applications without separate models
- Large variant requires ~1.3GB weights and significant VRAM for fine-tuning
- Not updated to incorporate newer Portuguese web data
- Newer multilingual models like mDeBERTa may be competitive on some tasks
Tags
transformerspytorchjaxbertfill-maskptdataset:brWaClicense:mitendpoints_compatibleregion:usdeploy:azure