From the model card
Fields below are copied from the tags and counters on the HuggingFace repository pyannote/speaker-diarization-community-1 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- pyannote
- Pipeline tag
- automatic-speech-recognition
- Library
- pyannote.audio
- License tag
cc-by-4.0— read the license file in the repo before relying on it- Papers cited
- arXiv:2104.03603, arXiv:2111.14448, arXiv:2012.01477, arXiv:2110.07058
- Downloads (HF counter at last fetch)
- 4,936,674
- Likes (HF counter at last fetch)
- 1,346
- Model card
- https://huggingface.co/pyannote/speaker-diarization-community-1
Use cases
- Transcribing multi-speaker meetings with speaker attribution
- Podcast or interview processing to label who speaks when
- Pre-processing audio before speaker-attributed ASR
- Research on speaker segmentation without gated model access
Pros
- Community model removes gated-access requirement of official pyannote models
- Integrates into pyannote pipeline chains
- MIT licensed
- Covers the full diarization pipeline in one call
Cons
- Lower accuracy than official pyannote models on diarization benchmarks
- Performance degrades with overlapping speech or more than 4 speakers
- Requires pyannote.audio with correct version pinning
- No detailed DER benchmark numbers published for this specific model
Tags
pyannote-audiopyannotepyannote-audio-pipelineaudiovoicespeechspeakerspeaker-diarizationspeaker-change-detectionvoice-activity-detectionoverlapped-speech-detectionautomatic-speech-recognitionarxiv:2104.03603arxiv:2111.14448arxiv:2012.01477arxiv:2110.07058license:cc-by-4.0region:us