AI Tools.

Search

automatic speech recognition by pyannote

voice-activity-detection

A pretrained voice activity detection pipeline from pyannote.audio, identifying speech segments in audio streams. It is trained on AMI, DIHARD, and VoxConverse corpora and outputs timestamped speech/non-speech labels.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository pyannote/voice-activity-detection at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
pyannote
Pipeline tag
automatic-speech-recognition
Library
pyannote.audio
License tag
mit — read the license file in the repo before relying on it
Datasets declared
ami, dihard, voxconverse
Downloads (HF counter at last fetch)
4,403,436
Likes (HF counter at last fetch)
241
Model card
https://huggingface.co/pyannote/voice-activity-detection

Use cases

  • Pre-processing audio before passing to ASR models
  • Filtering silence from podcast or meeting recordings
  • Building speaker turn detection pipelines
  • Reducing compute by skipping non-speech frames in streaming ASR

Pros

  • Trained on diverse real-world meeting and broadcast corpora
  • Outputs precise start/end timestamps, not just binary labels
  • Integrates directly into pyannote pipeline chains
  • MIT licensed with no restrictions on commercial use

Cons

  • Requires accepting pyannote's gated model terms on HuggingFace
  • Performance degrades on noisy environments like street audio
  • Not end-to-end — needs pyannote.audio installed with correct version
  • CPU inference is slow for real-time streaming applications

Tags

pyannote-audiopyannotepyannote-audio-pipelineaudiovoicespeechspeakervoice-activity-detectionautomatic-speech-recognitiondataset:amidataset:diharddataset:voxconverselicense:mitregion:us