AI Tools.

Search

fill mask by FacebookAI

roberta-large

RoBERTa large, the 355M-parameter version of Facebook AI's strongly trained BERT variant, offering doubled hidden size and additional attention heads over RoBERTa base. It provides stronger NLU accuracy at roughly 4x the inference compute cost of the base variant. Used where task accuracy on complex English language understanding outweighs latency constraints.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository FacebookAI/roberta-large at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
FacebookAI
Pipeline tag
fill-mask
Library
Transformers
Framework tags
PyTorch, TensorFlow, JAX
Weight formats
ONNX, safetensors
License tag
mit — read the license file in the repo before relying on it
Language tags
English (en)
Papers cited
arXiv:1907.11692, arXiv:1806.02847
Datasets declared
bookcorpus, wikipedia
Downloads (HF counter at last fetch)
6,546,673
Likes (HF counter at last fetch)
319
Model card
https://huggingface.co/FacebookAI/roberta-large

Use cases

  • High-accuracy text classification where inference latency is not critical
  • NLI and complex reasoning tasks requiring strong language understanding
  • Extractive QA on dense or technical passages
  • Research baseline for NLU benchmarks requiring a strong encoder
  • High-quality sentence embedding when lighter models underperform

Pros

  • Strong NLU performance from more parameters plus strong RoBERTa training
  • Multi-framework support (PyTorch, TF, JAX, ONNX, safetensors)
  • MIT license; widely published benchmark results for straightforward comparison
  • Dynamic masking pre-training generalizes better than static BERT masking

Cons

  • ~4x inference cost vs. RoBERTa base for marginal gains on simpler tasks
  • English-only; 512-token context limit
  • Encoder-only — cannot generate text
  • Surpassed by DeBERTa-v3-large and other newer encoders on most NLU benchmarks
  • High memory footprint limits use in latency-sensitive or edge deployments

Tags

transformerspytorchtfjaxonnxsafetensorsrobertafill-maskexbertendataset:bookcorpusdataset:wikipediaarxiv:1907.11692arxiv:1806.02847license:mitendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure