From the model card
Fields below are copied from the tags and counters on the HuggingFace repository pyannote/segmentation-3.0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- pyannote
- Pipeline tag
- voice-activity-detection
- Library
- pyannote.audio
- Framework tags
- PyTorch
- License tag
mit— read the license file in the repo before relying on it- Downloads (HF counter at last fetch)
- 5,794,665
- Likes (HF counter at last fetch)
- 1,653
- Model card
- https://huggingface.co/pyannote/segmentation-3.0
Use cases
- Voice activity detection to identify speech vs. non-speech regions
- Speaker change detection as preprocessing for downstream diarization
- Overlapping speech detection in multi-party conversations
- Audio preprocessing to remove silence before ASR
- Component in pyannote diarization pipeline
Pros
- MIT license
- Handles voice activity, speaker change, and overlapping speech in a single model
- Can run standalone for VAD without the full diarization stack
- State-of-the-art segmentation performance on pyannote benchmarks
- Integrates directly with speaker-diarization-3.1
Cons
- Requires HuggingFace token acceptance for download
- Frame-level model output requires post-processing for usable timestamps
- Overlapping speech detection accuracy degrades with more than 2 simultaneous speakers
- Not designed for keyword spotting or speech content analysis
- Performance varies with recording quality and background noise level
Tags
pyannote-audiopytorchpyannotepyannote-audio-modelaudiovoicespeechspeakerspeaker-diarizationspeaker-change-detectionspeaker-segmentationvoice-activity-detectionoverlapped-speech-detectionresegmentationlicense:mitregion:us