AI Tools.

Search

fill mask by neuralmind

bert-large-portuguese-cased

BERTimbau-large is a Portuguese BERT-large model pretrained from scratch on a 2.7B-word Portuguese corpus. It provides strong contextual representations for Brazilian and European Portuguese NLP tasks.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository neuralmind/bert-large-portuguese-cased at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
neuralmind
Pipeline tag
fill-mask
Library
Transformers
Framework tags
PyTorch, JAX
License tag
mit — read the license file in the repo before relying on it
Language tags
Portuguese (pt)
Datasets declared
brWaC
Downloads (HF counter at last fetch)
1,561,757
Likes (HF counter at last fetch)
74
Model card
https://huggingface.co/neuralmind/bert-large-portuguese-cased

Use cases

  • Portuguese text classification and entity recognition
  • Sentiment analysis on Portuguese social media and news
  • Portuguese question answering and reading comprehension
  • Fine-tuning base for Brazilian Portuguese NLP applications

Pros

  • Pretrained on Portuguese text — far outperforms multilingual BERT on Portuguese tasks
  • Apache-2.0 licensed
  • Cased variant preserves proper noun information
  • Well-benchmarked on standard Portuguese NLP datasets

Cons

  • Portuguese-only — not useful for bilingual applications without separate models
  • Large variant requires ~1.3GB weights and significant VRAM for fine-tuning
  • Not updated to incorporate newer Portuguese web data
  • Newer multilingual models like mDeBERTa may be competitive on some tasks

Tags

transformerspytorchjaxbertfill-maskptdataset:brWaClicense:mitendpoints_compatibleregion:usdeploy:azure