From the model card
Fields below are copied from the tags and counters on the HuggingFace repository openai/whisper-large-v3 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- openai
- Pipeline tag
- automatic-speech-recognition
- Library
- Transformers
- Framework tags
- PyTorch, JAX
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- English (en), Chinese (zh), German (de), Spanish (es), Russian (ru), Korean (ko), French (fr), Japanese (ja), Portuguese (pt), Turkish (tr), Polish (pl), Catalan (ca), Dutch (nl), Arabic (ar), Swedish (sv), Italian (it), Indonesian (id), Hindi (hi), Finnish (fi), Vietnamese (vi), Hebrew (he), Ukrainian (uk), Greek (el), Malay (ms), Czech (cs), Romanian (ro), Danish (da), Hungarian (hu), Tamil (ta), Norwegian (no), Thai (th), Urdu (ur), Croatian (hr), Bulgarian (bg), Lithuanian (lt), Latin (la), Māori (mi), Malayalam (ml), Welsh (cy), Slovak (sk), Telugu (te), Persian (fa), Latvian (lv), Bangla (bn), Serbian (sr), Azerbaijani (az), Slovenian (sl), Kannada (kn), Estonian (et), Macedonian (mk), Breton (br), Basque (eu), Icelandic (is), Armenian (hy), Nepali (ne), Mongolian (mn), Bosnian (bs), Kazakh (kk), Albanian (sq), Swahili (sw), Galician (gl), Marathi (mr), Punjabi (pa), Sinhala (si), Khmer (km), Shona (sn), Yoruba (yo), Somali (so), Afrikaans (af), Occitan (oc), Georgian (ka), Belarusian (be), Tajik (tg), Sindhi (sd), Gujarati (gu), Amharic (am), Yiddish (yi), Lao (lo), Uzbek (uz), Faroese (fo), Haitian Creole (ht), Pashto (ps), Turkmen (tk), Norwegian Nynorsk (nn), Maltese (mt), Sanskrit (sa), Luxembourgish (lb), Burmese (my), Tibetan (bo), Filipino (tl), Malagasy (mg), Assamese (as), Tatar (tt), Hawaiian (haw), Lingala (ln), Hausa (ha), Bashkir (ba), Javanese (jw), Sundanese (su)
- Papers cited
- arXiv:2212.04356
- Downloads (HF counter at last fetch)
- 4,908,520
- Likes (HF counter at last fetch)
- 6,232
- Model card
- https://huggingface.co/openai/whisper-large-v3
Use cases
- High-accuracy multilingual transcription where quality takes precedence over speed
- Long-form audio transcription (lectures, interviews, documentaries)
- Low-resource language transcription where smaller models underperform
- ASR research baseline requiring the best available open-weight transcription quality
- Subtitle generation for multilingual video content
Pros
- Apache 2.0 license for unrestricted commercial use
- 99+ language support at top-tier open-weight transcription quality
- Standard HuggingFace Transformers integration
- Benchmark-leading accuracy across multiple language ASR evaluations
Cons
- High GPU compute requirements — realtime transcription on long audio needs A100-class hardware
- Transcription latency on CPU is impractical for real-time use
- Large-v3-Turbo provides similar quality at lower cost for most use cases
- Word-level timestamps require additional inference passes or post-processing
- Diarization requires external combination with pyannote
Tags
transformerspytorchjaxsafetensorswhisperautomatic-speech-recognitionaudiohf-asr-leaderboardenzhdeesrukofrjapttrplca