AI Tools.

Search

feature extraction by facebook

w2v-bert-2.0

Meta's wav2vec-BERT 2.0 is a self-supervised speech encoder that combines contrastive learning with masked language modeling objectives. It serves as the backbone for Seamless and other Meta speech recognition and translation systems.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository facebook/w2v-bert-2.0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
facebook
Pipeline tag
feature-extraction
Library
Transformers
Weight formats
safetensors
License tag
mit — read the license file in the repo before relying on it
Language tags
Afrikaans (af), Amharic (am), Arabic (ar), Assamese (as), Azerbaijani (az), Belarusian (be), Bangla (bn), Bosnian (bs), Bulgarian (bg), Catalan (ca), Czech (cs), Chinese (zh), Welsh (cy), Danish (da), German (de), Greek (el), English (en), Estonian (et), Finnish (fi), French (fr), Odia (or), Oromo (om), Irish (ga), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Igbo (ig), Indonesian (id), Icelandic (is), Italian (it), Javanese (jv), Japanese (ja), Kannada (kn), Georgian (ka), Kazakh (kk), Mongolian (mn), Khmer (km), Kyrgyz (ky), Korean (ko), Lao (lo), Lingala (ln), Lithuanian (lt), Luxembourgish (lb), Ganda (lg), Latvian (lv), Malayalam (ml), Marathi (mr), Macedonian (mk), Maltese (mt), Māori (mi), Burmese (my), Dutch (nl), Norwegian Bokmål (nb), Nepali (ne), Nyanja (ny), Occitan (oc), Punjabi (pa), Pashto (ps), Persian (fa), Polish (pl), Portuguese (pt), Romanian (ro), Russian (ru), Slovak (sk), Slovenian (sl), Shona (sn), Sindhi (sd), Somali (so), Spanish (es), Serbian (sr), Swedish (sv), Swahili (sw), Tamil (ta), Telugu (te), Tajik (tg), Filipino (tl), Thai (th), Turkish (tr), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Wolof (wo), Xhosa (xh), Yoruba (yo), Malay (ms), Zulu (zu), Moroccan Arabic (ary), Egyptian Arabic (arz), Cantonese (yue), Kabuverdianu (kea)
Papers cited
arXiv:2312.05187
Downloads (HF counter at last fetch)
2,245,550
Likes (HF counter at last fetch)
227
Model card
https://huggingface.co/facebook/w2v-bert-2.0

Use cases

  • Pre-training foundation for downstream ASR fine-tuning
  • Multilingual speech representation learning
  • Feature extraction for speech classification tasks
  • Building the encoder stage in speech-to-text translation pipelines

Pros

  • Trained on 4.5M hours of unlabeled speech across 143 languages
  • Achieves strong ASR results with limited labeled data when fine-tuned
  • MIT licensed
  • Backbone of Meta's production Seamless translation models

Cons

  • Requires fine-tuning on labeled data before producing transcripts
  • Large model size makes edge deployment impractical
  • Training from scratch requires massive compute
  • Not plug-and-play — needs connectionist temporal classification head for ASR

Tags

transformerssafetensorswav2vec2-bertfeature-extractionafamarasazbebnbsbgcacszhcydadeel