From the model card
Fields below are copied from the tags and counters on the HuggingFace repository sentence-transformers/all-distilroberta-v1 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- sentence-transformers
- Pipeline tag
- sentence-similarity
- Library
- Sentence Transformers, Transformers
- Framework tags
- PyTorch, Rust (candle)
- Weight formats
- ONNX, safetensors, OpenVINO
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:1904.06472, arXiv:2102.07033, arXiv:2104.08727, arXiv:1704.05179, arXiv:1810.09305
- Datasets declared
- s2orc, flax-sentence-embeddings/stackexchange_xml, ms_marco, gooaq, yahoo_answers_topics, code_search_net, search_qa, eli5 and 13 more on the model card
- Downloads (HF counter at last fetch)
- 2,469,836
- Likes (HF counter at last fetch)
- 43
- Model card
- https://huggingface.co/sentence-transformers/all-distilroberta-v1
Use cases
- General semantic textual similarity
- Semantic clustering of medium-length English text
- Baseline comparison against larger sentence-transformer models
- Retrieval applications where 768-dim representation quality is needed
Pros
- 768-dim output vs 384-dim for MiniLM — higher-capacity representations
- Apache-2.0 licensed
- Trained on diverse billion-sentence dataset
- DistilRoBERTa backbone runs faster than full RoBERTa
Cons
- Outperformed by all-mpnet-base-v2 at similar compute cost
- English-only
- Newer GTE and E5 models significantly outperform on MTEB benchmarks
- Not optimized for asymmetric retrieval tasks
Tags
sentence-transformerspytorchrustonnxsafetensorsopenvinorobertafill-maskfeature-extractionsentence-similaritytransformersendataset:s2orcdataset:flax-sentence-embeddings/stackexchange_xmldataset:ms_marcodataset:gooaqdataset:yahoo_answers_topicsdataset:code_search_netdataset:search_qadataset:eli5