AI Tools.

Search

text to speech models

7 models · ranked by HuggingFace downloads

Kokoro-82M

by hexgrad

Kokoro-82M is a compact 82-million-parameter text-to-speech model fine-tuned from StyleTTS2, targeting natural-sounding English speech synthesis at a size runnable on CPU or modest GPU. Released under Apache 2.0 with a HuggingFace DOI, it gained attention as a high-quality open TTS model at significantly smaller scale than most alternatives. It supports multiple English voice styles.

11,290,774 ↓ · 6,806 ♡

XTTS-v2

by coqui

XTTS-v2 is Coqui's multilingual text-to-speech model supporting 17 languages with voice cloning from a short audio sample. It uses a GPT-style decoder for speech token generation, enabling zero-shot speaker cloning without fine-tuning. The model was released before Coqui's closure and remains available under a non-standard license.

7,415,626 ↓ · 3,762 ♡

chatterbox

by ResembleAI

Chatterbox is Resemble AI's open-source text-to-speech model offering voice cloning and expressive speech synthesis. It is designed as a production-grade TTS system with controllable prosody and emotion.

1,696,137 ↓ · 1,777 ♡

OmniVoice

by k2-fsa

OmniVoice from k2-fsa is a multilingual speech model targeting end-to-end ASR and voice processing tasks. Published as part of the k2/Lhotse/sherpa-onnx ecosystem for server and edge speech applications.

1,013,539 ↓ · 1,334 ♡

s2-pro

by fishaudio

s2-pro is Fish Audio's multilingual text-to-speech model supporting over 80 languages with instruction-following capabilities, described in arXiv:2603.08823. It is designed for zero-shot voice cloning and cross-lingual synthesis by conditioning on speaker reference audio and natural language prompts. The license is marked 'other', meaning specific usage restrictions apply beyond standard open-source terms.

414,738 ↓ · 1,280 ♡

Qwen3-TTS-12Hz-1.7B-VoiceDesign

by Qwen

Qwen3-TTS VoiceDesign is a 1.7B text-to-speech model operating at 12Hz token rate, designed to support custom voice creation alongside standard TTS. It covers multiple languages and generates expressive speech from text input. Apache-2.0 licensed and part of Qwen's audio model family.

404,061 ↓ · 392 ♡

VoxCPM2

by openbmb

VoxCPM2 is a multilingual text-to-speech model from OpenBMB supporting over 35 languages, with explicit voice-cloning and voice-design capabilities built on a diffusion-based audio synthesis approach. It covers a wide geographic range including East Asian, Southeast Asian, European, and Middle Eastern languages. The model is released under Apache-2.0.

402,878 ↓ · 1,539 ♡