From the model card
Fields below are copied from the tags and counters on the HuggingFace repository hexgrad/Kokoro-82M at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- hexgrad
- Pipeline tag
- text-to-speech
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model yl4579/StyleTTS2-LJSpeech
- fine-tune of yl4579/StyleTTS2-LJSpeech
- Language tags
- English (en)
- Papers cited
- arXiv:2306.07691, arXiv:2203.02395
- Downloads (HF counter at last fetch)
- 11,290,774
- Likes (HF counter at last fetch)
- 6,806
- Model card
- https://huggingface.co/hexgrad/Kokoro-82M
Use cases
- Local TTS for accessibility tools and screen readers without API cost
- Podcast and audiobook content creation from text
- Voice assistant response generation on-device or in lightweight servers
- Narration generation for video content at low compute cost
- Research into efficient TTS at sub-100M parameter scale
Pros
- Apache 2.0 license for unrestricted commercial use
- 82M parameters enables CPU and low-end GPU inference
- Natural prosody quality for its parameter count, based on StyleTTS2
- Multiple English voice styles available from a single checkpoint
Cons
- English-only; no multilingual TTS capability
- Prosody and naturalness below larger TTS models for demanding audiobook production
- Limited control over speaking rate and emphasis compared to larger commercial TTS APIs
- Community model without a major lab's production testing or SLA
- Fine-tuning requires StyleTTS2 training expertise
Tags
text-to-speechenarxiv:2306.07691arxiv:2203.02395base_model:yl4579/StyleTTS2-LJSpeechbase_model:finetune:yl4579/StyleTTS2-LJSpeechdoi:10.57967/hf/4329license:apache-2.0region:us