AI Tools.

Search

automatic speech recognition by MahmoudAshraf

mms-300m-1130-forced-aligner

MMS-300M-1130-forced-aligner is Meta's 300M parameter wav2vec2-based model fine-tuned for forced phoneme-level alignment across 1,130 languages. It takes audio and a text transcript as input and outputs word- or phoneme-level timestamps, enabling subtitle synchronization and linguistic documentation at scale. The CC-BY-NC-4.0 license restricts commercial deployment.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository MahmoudAshraf/mms-300m-1130-forced-aligner at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
MahmoudAshraf
Pipeline tag
automatic-speech-recognition
Library
Transformers
Framework tags
PyTorch
Weight formats
safetensors
License tag
cc-by-nc-4.0 — read the license file in the repo before relying on it
Language tags
Abkhazian (ab), Afrikaans (af), Akan (ak), Amharic (am), Arabic (ar), Assamese (as), Avaric (av), Aymara (ay), Azerbaijani (az), Bashkir (ba), Bambara (bm), Belarusian (be), Bangla (bn), Bislama (bi), Tibetan (bo), Serbian (Latin) (sh), Breton (br), Bulgarian (bg), Catalan (ca), Czech (cs), Chechen (ce), Chuvash (cv), Kurdish (ku), Welsh (cy), Danish (da), German (de), Divehi (dv), Dzongkha (dz), Greek (el), English (en), Esperanto (eo), Estonian (et), Basque (eu), Ewe (ee), Faroese (fo), Persian (fa), Fijian (fj), Finnish (fi), French (fr), Western Frisian (fy), Fula (ff), Irish (ga), Galician (gl), Guarani (gn), Gujarati (gu), Chinese (zh), Haitian Creole (ht), Hausa (ha), Hebrew (he), Hindi (hi), Hungarian (hu), Armenian (hy), Igbo (ig), Interlingua (ia), Malay (ms), Icelandic (is), Italian (it), Javanese (jv), Japanese (ja), Kannada (kn), Georgian (ka), Kazakh (kk), Kanuri (kr), Khmer (km), Kikuyu (ki), Kinyarwanda (rw), Kyrgyz (ky), Korean (ko), Komi (kv), Lao (lo), Latin (la), Latvian (lv), Lingala (ln), Lithuanian (lt), Luxembourgish (lb), Ganda (lg), Marshallese (mh), Malayalam (ml), Marathi (mr), Macedonian (mk), Malagasy (mg), Maltese (mt), Mongolian (mn), Māori (mi), Burmese (my), Dutch (nl), Norwegian (no), Nepali (ne), Nyanja (ny), Occitan (oc), Oromo (om), Odia (or), Ossetic (os), Punjabi (pa), Polish (pl), Portuguese (pt), Pashto (ps), Quechua (qu), Romanian (ro), Rundi (rn), Russian (ru), Sango (sg), Slovak (sk), Slovenian (sl), Samoan (sm), Shona (sn), Sindhi (sd), Somali (so), Spanish (es), Albanian (sq), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Tatar (tt), Telugu (te), Tajik (tg), Filipino (tl), Thai (th), Tigrinya (ti), Tsonga (ts), Turkish (tr), Ukrainian (uk), Vietnamese (vi), Wolof (wo), Xhosa (xh), Yoruba (yo), Zulu (zu), Zhuang (za)
Downloads (HF counter at last fetch)
2,430,219
Likes (HF counter at last fetch)
102
Model card
https://huggingface.co/MahmoudAshraf/mms-300m-1130-forced-aligner

Use cases

  • Automated subtitle timestamp generation from existing transcripts
  • Phoneme-level alignment for low-resource language documentation
  • Speech data annotation for multilingual TTS training corpus creation
  • Linguistic research on timing patterns across diverse language families

Pros

  • Supports 1,130 languages, far exceeding other forced alignment tools
  • Produces fine-grained word and phoneme-level timestamps
  • wav2vec2 backbone integrates directly with HuggingFace ecosystem tooling

Cons

  • CC-BY-NC-4.0 license prohibits commercial deployment
  • Requires a pre-existing text transcript as input — not a standalone ASR model
  • Accuracy drops significantly on noisy or heavily accented audio recordings

Tags

transformerspytorchsafetensorswav2vec2automatic-speech-recognitionmmsaudiovoicespeechforced-alignmentabafakamarasavayazba