From the model card
Fields below are copied from the tags and counters on the HuggingFace repository openbmb/VoxCPM2 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- openbmb
- Pipeline tag
- text-to-speech
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual; Chinese (zh), English (en), Arabic (ar), Burmese (my), Danish (da), Dutch (nl), Finnish (fi), French (fr), German (de), Greek (el), Hebrew (he), Hindi (hi), Indonesian (id), Italian (it), Japanese (ja), Khmer (km), Korean (ko), Lao (lo), Malay (ms), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Spanish (es), Swahili (sw), Swedish (sv), Filipino (tl), Thai (th), Turkish (tr), Vietnamese (vi)
- Papers cited
- arXiv:2509.24650
- Downloads (HF counter at last fetch)
- 402,878
- Likes (HF counter at last fetch)
- 1,539
- Model card
- https://huggingface.co/openbmb/VoxCPM2
Use cases
- Multilingual audiobook and podcast narration across 35+ languages
- Voice cloning for content localization into low-resource languages
- Building TTS pipelines for Southeast Asian language applications
- Voice design research exploring controllable speaker characteristics
- Accessibility tooling requiring high-coverage multilingual speech synthesis
Pros
- 35+ language coverage is among the broadest of openly available TTS models
- Voice cloning support enables speaker-adapted synthesis without retraining
- Diffusion-based synthesis typically produces more natural prosody than autoregressive alternatives
- Apache-2.0 license allows unrestricted commercial and research use
- Strong community traction with 1,434 likes and nearly 585K downloads
Cons
- Diffusion-based inference is slower than streaming autoregressive TTS models
- Voice quality for lower-resource languages (e.g., Burmese, Khmer) may lag behind high-resource ones
- No native streaming/chunked synthesis API described, limiting real-time use cases
- Model size and compute requirements are not prominently documented
- Voice cloning quality is sensitive to reference audio length and recording conditions
Tags
voxcpmsafetensorstext-to-speechttsmultilingualvoice-cloningvoice-designdiffusionaudiozhenarmydanlfifrdeelhe