From the model card
Fields below are copied from the tags and counters on the HuggingFace repository pyannote/voice-activity-detection at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- pyannote
- Pipeline tag
- automatic-speech-recognition
- Library
- pyannote.audio
- License tag
mit— read the license file in the repo before relying on it- Datasets declared
- ami, dihard, voxconverse
- Downloads (HF counter at last fetch)
- 4,403,436
- Likes (HF counter at last fetch)
- 241
- Model card
- https://huggingface.co/pyannote/voice-activity-detection
Use cases
- Pre-processing audio before passing to ASR models
- Filtering silence from podcast or meeting recordings
- Building speaker turn detection pipelines
- Reducing compute by skipping non-speech frames in streaming ASR
Pros
- Trained on diverse real-world meeting and broadcast corpora
- Outputs precise start/end timestamps, not just binary labels
- Integrates directly into pyannote pipeline chains
- MIT licensed with no restrictions on commercial use
Cons
- Requires accepting pyannote's gated model terms on HuggingFace
- Performance degrades on noisy environments like street audio
- Not end-to-end — needs pyannote.audio installed with correct version
- CPU inference is slow for real-time streaming applications
Tags
pyannote-audiopyannotepyannote-audio-pipelineaudiovoicespeechspeakervoice-activity-detectionautomatic-speech-recognitiondataset:amidataset:diharddataset:voxconverselicense:mitregion:us