From the model card
Fields below are copied from the tags and counters on the HuggingFace repository ResembleAI/chatterbox at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- ResembleAI
- Pipeline tag
- text-to-speech
- License tag
mit— read the license file in the repo before relying on it- Language tags
- Arabic (ar), Danish (da), German (de), Greek (el), English (en), Spanish (es), Finnish (fi), French (fr), Hebrew (he), Hindi (hi), Italian (it), Japanese (ja), Korean (ko), Malay (ms), Dutch (nl), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Swedish (sv), Swahili (sw), Turkish (tr), Chinese (zh)
- Downloads (HF counter at last fetch)
- 1,696,137
- Likes (HF counter at last fetch)
- 1,777
- Model card
- https://huggingface.co/ResembleAI/chatterbox
Use cases
- Voice cloning from short reference audio samples
- Expressive narration for audiobooks and podcasts
- Custom voice assistant deployment
- Generating diverse synthetic voices for training data
Pros
- Voice cloning with short reference audio — no lengthy enrollment
- Open-source release with Apache-2.0 license
- Controllable expressiveness beyond flat TTS systems
- Maintained by Resemble AI with commercial-grade quality targets
Cons
- Voice cloning raises consent and misuse concerns — deploy responsibly
- Quality depends heavily on reference audio cleanliness
- Inference speed may not meet real-time requirements on CPU
- Limited language support compared to Coqui or Kokoro TTS models
Tags
chatterboxtext-to-speechspeechspeech-generationvoice-cloningmultilingual-ttsardadeelenesfifrhehiitjakoms