by hexgrad
Kokoro-82M is a compact 82-million-parameter text-to-speech model fine-tuned from StyleTTS2, targeting natural-sounding English speech synthesis at a size runnable on CPU or modest GPU. Released under Apache 2.0 with a HuggingFace DOI, it gained attention as a high-quality open TTS model at significantly smaller scale than most alternatives. It supports multiple English voice styles.
11,290,774 ↓ · 6,806 ♡
by coqui
XTTS-v2 is Coqui's multilingual text-to-speech model supporting 17 languages with voice cloning from a short audio sample. It uses a GPT-style decoder for speech token generation, enabling zero-shot speaker cloning without fine-tuning. The model was released before Coqui's closure and remains available under a non-standard license.
7,415,626 ↓ · 3,762 ♡
by ResembleAI
Chatterbox is Resemble AI's open-source text-to-speech model offering voice cloning and expressive speech synthesis. It is designed as a production-grade TTS system with controllable prosody and emotion.
1,696,137 ↓ · 1,777 ♡
by k2-fsa
OmniVoice from k2-fsa is a multilingual speech model targeting end-to-end ASR and voice processing tasks. Published as part of the k2/Lhotse/sherpa-onnx ecosystem for server and edge speech applications.
1,013,539 ↓ · 1,334 ♡
by fishaudio
s2-pro is Fish Audio's multilingual text-to-speech model supporting over 80 languages with instruction-following capabilities, described in arXiv:2603.08823. It is designed for zero-shot voice cloning and cross-lingual synthesis by conditioning on speaker reference audio and natural language prompts. The license is marked 'other', meaning specific usage restrictions apply beyond standard open-source terms.
414,738 ↓ · 1,280 ♡
by Qwen
Qwen3-TTS VoiceDesign is a 1.7B text-to-speech model operating at 12Hz token rate, designed to support custom voice creation alongside standard TTS. It covers multiple languages and generates expressive speech from text input. Apache-2.0 licensed and part of Qwen's audio model family.
404,061 ↓ · 392 ♡
by openbmb
VoxCPM2 is a multilingual text-to-speech model from OpenBMB supporting over 35 languages, with explicit voice-cloning and voice-design capabilities built on a diffusion-based audio synthesis approach. It covers a wide geographic range including East Asian, Southeast Asian, European, and Middle Eastern languages. The model is released under Apache-2.0.
402,878 ↓ · 1,539 ♡