AI Tools.

Search

automatic speech recognition by openai

whisper-large-v3

Whisper Large-v3 is OpenAI's full-size ASR model supporting 99+ languages, trained on 680,000 hours of multilingual audio. It delivers state-of-the-art transcription accuracy across languages at the cost of significant inference compute. Apache 2.0 licensed. The Large-v3-Turbo variant (a distilled version) provides similar quality at lower cost for most use cases.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository openai/whisper-large-v3 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
openai
Pipeline tag
automatic-speech-recognition
Library
Transformers
Framework tags
PyTorch, JAX
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
English (en), Chinese (zh), German (de), Spanish (es), Russian (ru), Korean (ko), French (fr), Japanese (ja), Portuguese (pt), Turkish (tr), Polish (pl), Catalan (ca), Dutch (nl), Arabic (ar), Swedish (sv), Italian (it), Indonesian (id), Hindi (hi), Finnish (fi), Vietnamese (vi), Hebrew (he), Ukrainian (uk), Greek (el), Malay (ms), Czech (cs), Romanian (ro), Danish (da), Hungarian (hu), Tamil (ta), Norwegian (no), Thai (th), Urdu (ur), Croatian (hr), Bulgarian (bg), Lithuanian (lt), Latin (la), Māori (mi), Malayalam (ml), Welsh (cy), Slovak (sk), Telugu (te), Persian (fa), Latvian (lv), Bangla (bn), Serbian (sr), Azerbaijani (az), Slovenian (sl), Kannada (kn), Estonian (et), Macedonian (mk), Breton (br), Basque (eu), Icelandic (is), Armenian (hy), Nepali (ne), Mongolian (mn), Bosnian (bs), Kazakh (kk), Albanian (sq), Swahili (sw), Galician (gl), Marathi (mr), Punjabi (pa), Sinhala (si), Khmer (km), Shona (sn), Yoruba (yo), Somali (so), Afrikaans (af), Occitan (oc), Georgian (ka), Belarusian (be), Tajik (tg), Sindhi (sd), Gujarati (gu), Amharic (am), Yiddish (yi), Lao (lo), Uzbek (uz), Faroese (fo), Haitian Creole (ht), Pashto (ps), Turkmen (tk), Norwegian Nynorsk (nn), Maltese (mt), Sanskrit (sa), Luxembourgish (lb), Burmese (my), Tibetan (bo), Filipino (tl), Malagasy (mg), Assamese (as), Tatar (tt), Hawaiian (haw), Lingala (ln), Hausa (ha), Bashkir (ba), Javanese (jw), Sundanese (su)
Papers cited
arXiv:2212.04356
Downloads (HF counter at last fetch)
4,908,520
Likes (HF counter at last fetch)
6,232
Model card
https://huggingface.co/openai/whisper-large-v3

Use cases

  • High-accuracy multilingual transcription where quality takes precedence over speed
  • Long-form audio transcription (lectures, interviews, documentaries)
  • Low-resource language transcription where smaller models underperform
  • ASR research baseline requiring the best available open-weight transcription quality
  • Subtitle generation for multilingual video content

Pros

  • Apache 2.0 license for unrestricted commercial use
  • 99+ language support at top-tier open-weight transcription quality
  • Standard HuggingFace Transformers integration
  • Benchmark-leading accuracy across multiple language ASR evaluations

Cons

  • High GPU compute requirements — realtime transcription on long audio needs A100-class hardware
  • Transcription latency on CPU is impractical for real-time use
  • Large-v3-Turbo provides similar quality at lower cost for most use cases
  • Word-level timestamps require additional inference passes or post-processing
  • Diarization requires external combination with pyannote

Tags

transformerspytorchjaxsafetensorswhisperautomatic-speech-recognitionaudiohf-asr-leaderboardenzhdeesrukofrjapttrplca