AI Tools.

Search

automatic speech recognition by pyannote

speaker-diarization-community-1

A community-supported speaker diarization pipeline from pyannote.audio that segments multi-speaker audio into per-speaker turns. It combines voice activity detection, speaker embedding, and clustering steps into a single callable pipeline.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository pyannote/speaker-diarization-community-1 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
pyannote
Pipeline tag
automatic-speech-recognition
Library
pyannote.audio
License tag
cc-by-4.0 — read the license file in the repo before relying on it
Papers cited
arXiv:2104.03603, arXiv:2111.14448, arXiv:2012.01477, arXiv:2110.07058
Downloads (HF counter at last fetch)
4,936,674
Likes (HF counter at last fetch)
1,346
Model card
https://huggingface.co/pyannote/speaker-diarization-community-1

Use cases

  • Transcribing multi-speaker meetings with speaker attribution
  • Podcast or interview processing to label who speaks when
  • Pre-processing audio before speaker-attributed ASR
  • Research on speaker segmentation without gated model access

Pros

  • Community model removes gated-access requirement of official pyannote models
  • Integrates into pyannote pipeline chains
  • MIT licensed
  • Covers the full diarization pipeline in one call

Cons

  • Lower accuracy than official pyannote models on diarization benchmarks
  • Performance degrades with overlapping speech or more than 4 speakers
  • Requires pyannote.audio with correct version pinning
  • No detailed DER benchmark numbers published for this specific model

Tags

pyannote-audiopyannotepyannote-audio-pipelineaudiovoicespeechspeakerspeaker-diarizationspeaker-change-detectionvoice-activity-detectionoverlapped-speech-detectionautomatic-speech-recognitionarxiv:2104.03603arxiv:2111.14448arxiv:2012.01477arxiv:2110.07058license:cc-by-4.0region:us