From the model card
Fields below are copied from the tags and counters on the HuggingFace repository FacebookAI/roberta-large at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- FacebookAI
- Pipeline tag
- fill-mask
- Library
- Transformers
- Framework tags
- PyTorch, TensorFlow, JAX
- Weight formats
- ONNX, safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:1907.11692, arXiv:1806.02847
- Datasets declared
- bookcorpus, wikipedia
- Downloads (HF counter at last fetch)
- 6,546,673
- Likes (HF counter at last fetch)
- 319
- Model card
- https://huggingface.co/FacebookAI/roberta-large
Use cases
- High-accuracy text classification where inference latency is not critical
- NLI and complex reasoning tasks requiring strong language understanding
- Extractive QA on dense or technical passages
- Research baseline for NLU benchmarks requiring a strong encoder
- High-quality sentence embedding when lighter models underperform
Pros
- Strong NLU performance from more parameters plus strong RoBERTa training
- Multi-framework support (PyTorch, TF, JAX, ONNX, safetensors)
- MIT license; widely published benchmark results for straightforward comparison
- Dynamic masking pre-training generalizes better than static BERT masking
Cons
- ~4x inference cost vs. RoBERTa base for marginal gains on simpler tasks
- English-only; 512-token context limit
- Encoder-only — cannot generate text
- Surpassed by DeBERTa-v3-large and other newer encoders on most NLU benchmarks
- High memory footprint limits use in latency-sensitive or edge deployments
Tags
transformerspytorchtfjaxonnxsafetensorsrobertafill-maskexbertendataset:bookcorpusdataset:wikipediaarxiv:1907.11692arxiv:1806.02847license:mitendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure