From the model card
Fields below are copied from the tags and counters on the HuggingFace repository facebook/w2v-bert-2.0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Pipeline tag
- feature-extraction
- Library
- Transformers
- Weight formats
- safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bangla (bn), Bosnian (bs), Bulgarian (bg), Catalan (ca), Czech (cs), Chinese (zh), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Estonian (et), Finnish (fi), French (fr), Odia (or), Oromo (om), Irish (ga), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Igbo (ig), Indonesian (id), Icelandic (is), Italian (it), Javanese (jv), Japanese (ja), Kannada (kn), Georgian (ka), Kazakh (kk), Mongolian (mn), Khmer (km), Kyrgyz (ky), Korean (ko), Lao (lo), Lingala (ln), Lithuanian (lt), Luxembourgish (lb), Ganda (lg), Latvian (lv), Malayalam (ml), Marathi (mr), Macedonian (mk), Maltese (mt), Māori (mi), Burmese (my), Dutch (nl), Norwegian Bokmål (nb), Nepali (ne), Nyanja (ny), Occitan (oc), Punjabi (pa), Pashto (ps), Persian (fa), Polish (pl), Portuguese (pt), Romanian (ro), Russian (ru), Slovak (sk), Slovenian (sl), Shona (sn), Sindhi (sd), Somali (so), Spanish (es), Serbian (sr), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Tajik (tg), Filipino (tl), Thai (th), Turkish (tr), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Wolof (wo), Xhosa (xh), Yoruba (yo), Malay (ms), Zulu (zu), Moroccan Arabic (ary), Egyptian Arabic (arz), Cantonese (yue), Kabuverdianu (kea)
- Papers cited
- arXiv:2312.05187
- Downloads (HF counter at last fetch)
- 2,245,550
- Likes (HF counter at last fetch)
- 227
- Model card
- https://huggingface.co/facebook/w2v-bert-2.0
Use cases
- Pre-training foundation for downstream ASR fine-tuning
- Multilingual speech representation learning
- Feature extraction for speech classification tasks
- Building the encoder stage in speech-to-text translation pipelines
Pros
- Trained on 4.5M hours of unlabeled speech across 143 languages
- Achieves strong ASR results with limited labeled data when fine-tuned
- MIT licensed
- Backbone of Meta's production Seamless translation models
Cons
- Requires fine-tuning on labeled data before producing transcripts
- Large model size makes edge deployment impractical
- Training from scratch requires massive compute
- Not plug-and-play — needs connectionist temporal classification head for ASR
Tags
transformerssafetensorswav2vec2-bertfeature-extractionafamarasazbebnbsbgcacszhcydadeel