AI Tools.

Search

text to speech by coqui

XTTS-v2

XTTS-v2 is Coqui's multilingual text-to-speech model supporting 17 languages with voice cloning from a short audio sample. It uses a GPT-style decoder for speech token generation, enabling zero-shot speaker cloning without fine-tuning. The model was released before Coqui's closure and remains available under a non-standard license.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository coqui/XTTS-v2 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
coqui
Pipeline tag
text-to-speech
License tag
other — read the license file in the repo before relying on it
Downloads (HF counter at last fetch)
7,415,626
Likes (HF counter at last fetch)
3,762
Model card
https://huggingface.co/coqui/XTTS-v2

Use cases

  • Multilingual voice cloning for localization workflows
  • Zero-shot TTS from a 6-second speaker audio sample
  • Audiobook narration in supported languages
  • Game character voice generation with consistent speaker identity
  • Accessibility tools requiring personalized voice output

Pros

  • 17-language multilingual support including Portuguese, Polish, Turkish, and Arabic
  • Voice cloning from a short audio sample without fine-tuning
  • GPT-based decoder produces more natural prosody than older TTS models
  • Widely tested in the Coqui TTS open-source ecosystem

Cons

  • License is 'other' — not Apache/MIT; Coqui has closed operations, review terms carefully for commercial use
  • Voice cloning quality varies significantly with audio sample quality and duration
  • Inference requires more compute than simpler TTS architectures
  • No active maintenance following Coqui's closure
  • Output quality for low-resource languages in the 17-language set varies substantially

Tags

coquitext-to-speechlicense:otherregion:us