AI Tools.

Search

text to speech by openbmb

VoxCPM2

VoxCPM2 is a multilingual text-to-speech model from OpenBMB supporting over 35 languages, with explicit voice-cloning and voice-design capabilities built on a diffusion-based audio synthesis approach. It covers a wide geographic range including East Asian, Southeast Asian, European, and Middle Eastern languages. The model is released under Apache-2.0.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository openbmb/VoxCPM2 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
openbmb
Pipeline tag
text-to-speech
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual; Chinese (zh), English (en), Arabic (ar), Burmese (my), Danish (da), Dutch (nl), Finnish (fi), French (fr), German (de), Greek (el), Hebrew (he), Hindi (hi), Indonesian (id), Italian (it), Japanese (ja), Khmer (km), Korean (ko), Lao (lo), Malay (ms), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Spanish (es), Swahili (sw), Swedish (sv), Filipino (tl), Thai (th), Turkish (tr), Vietnamese (vi)
Papers cited
arXiv:2509.24650
Downloads (HF counter at last fetch)
402,878
Likes (HF counter at last fetch)
1,539
Model card
https://huggingface.co/openbmb/VoxCPM2

Use cases

  • Multilingual audiobook and podcast narration across 35+ languages
  • Voice cloning for content localization into low-resource languages
  • Building TTS pipelines for Southeast Asian language applications
  • Voice design research exploring controllable speaker characteristics
  • Accessibility tooling requiring high-coverage multilingual speech synthesis

Pros

  • 35+ language coverage is among the broadest of openly available TTS models
  • Voice cloning support enables speaker-adapted synthesis without retraining
  • Diffusion-based synthesis typically produces more natural prosody than autoregressive alternatives
  • Apache-2.0 license allows unrestricted commercial and research use
  • Strong community traction with 1,434 likes and nearly 585K downloads

Cons

  • Diffusion-based inference is slower than streaming autoregressive TTS models
  • Voice quality for lower-resource languages (e.g., Burmese, Khmer) may lag behind high-resource ones
  • No native streaming/chunked synthesis API described, limiting real-time use cases
  • Model size and compute requirements are not prominently documented
  • Voice cloning quality is sensitive to reference audio length and recording conditions

Tags

voxcpmsafetensorstext-to-speechttsmultilingualvoice-cloningvoice-designdiffusionaudiozhenarmydanlfifrdeelhe