From the model card
Fields below are copied from the tags and counters on the HuggingFace repository coqui/XTTS-v2 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- coqui
- Pipeline tag
- text-to-speech
- License tag
other— read the license file in the repo before relying on it- Downloads (HF counter at last fetch)
- 7,415,626
- Likes (HF counter at last fetch)
- 3,762
- Model card
- https://huggingface.co/coqui/XTTS-v2
Use cases
- Multilingual voice cloning for localization workflows
- Zero-shot TTS from a 6-second speaker audio sample
- Audiobook narration in supported languages
- Game character voice generation with consistent speaker identity
- Accessibility tools requiring personalized voice output
Pros
- 17-language multilingual support including Portuguese, Polish, Turkish, and Arabic
- Voice cloning from a short audio sample without fine-tuning
- GPT-based decoder produces more natural prosody than older TTS models
- Widely tested in the Coqui TTS open-source ecosystem
Cons
- License is 'other' — not Apache/MIT; Coqui has closed operations, review terms carefully for commercial use
- Voice cloning quality varies significantly with audio sample quality and duration
- Inference requires more compute than simpler TTS architectures
- No active maintenance following Coqui's closure
- Output quality for low-resource languages in the 17-language set varies substantially
Tags
coquitext-to-speechlicense:otherregion:us