From the model card
Fields below are copied from the tags and counters on the HuggingFace repository facebook/hubert-large-ls960-ft at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Pipeline tag
- automatic-speech-recognition
- Library
- Transformers
- Framework tags
- PyTorch, TensorFlow
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:2106.07447
- Datasets declared
- libri-light, librispeech_asr
- Downloads (HF counter at last fetch)
- 368,725
- Likes (HF counter at last fetch)
- 76
- Model card
- https://huggingface.co/facebook/hubert-large-ls960-ft
Use cases
- English speech-to-text transcription
- ASR research benchmark on LibriSpeech-style clean speech
- Foundation for domain-specific English ASR fine-tuning
- Audio feature extraction for downstream speech ML tasks
Pros
- Apache-2.0 license
- HuBERT pretraining provides strong general audio representations
- PyTorch and TF checkpoints available
- Published evaluation on LibriSpeech WER
Cons
- English-only — not suitable for multilingual ASR
- HuBERT-Large is large and slow compared to distilled Whisper variants
- No built-in punctuation — raw word sequences only
- Significantly outperformed by recent models on complex speech conditions (accents, noise)
Tags
transformerspytorchtfhubertautomatic-speech-recognitionspeechaudiohf-asr-leaderboardendataset:libri-lightdataset:librispeech_asrarxiv:2106.07447license:apache-2.0model-indexeval-resultsendpoints_compatibledeploy:azureregion:us