AI Tools.

Search

voice activity detection by pyannote

segmentation-3.0

Pyannote segmentation-3.0 is a speaker segmentation model for detecting speaker changes, overlapping speech, and voice activity in audio. It produces frame-level predictions used as input to the full speaker diarization pipeline. The model can also run standalone for voice activity detection or overlapped speech detection without the full diarization stack.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository pyannote/segmentation-3.0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
pyannote
Pipeline tag
voice-activity-detection
Library
pyannote.audio
Framework tags
PyTorch
License tag
mit — read the license file in the repo before relying on it
Downloads (HF counter at last fetch)
5,794,665
Likes (HF counter at last fetch)
1,653
Model card
https://huggingface.co/pyannote/segmentation-3.0

Use cases

  • Voice activity detection to identify speech vs. non-speech regions
  • Speaker change detection as preprocessing for downstream diarization
  • Overlapping speech detection in multi-party conversations
  • Audio preprocessing to remove silence before ASR
  • Component in pyannote diarization pipeline

Pros

  • MIT license
  • Handles voice activity, speaker change, and overlapping speech in a single model
  • Can run standalone for VAD without the full diarization stack
  • State-of-the-art segmentation performance on pyannote benchmarks
  • Integrates directly with speaker-diarization-3.1

Cons

  • Requires HuggingFace token acceptance for download
  • Frame-level model output requires post-processing for usable timestamps
  • Overlapping speech detection accuracy degrades with more than 2 simultaneous speakers
  • Not designed for keyword spotting or speech content analysis
  • Performance varies with recording quality and background noise level

Tags

pyannote-audiopytorchpyannotepyannote-audio-modelaudiovoicespeechspeakerspeaker-diarizationspeaker-change-detectionspeaker-segmentationvoice-activity-detectionoverlapped-speech-detectionresegmentationlicense:mitregion:us