From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- text-to-speech
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual
- Papers cited
- arXiv:2601.15621
- Downloads (HF counter at last fetch)
- 404,061
- Likes (HF counter at last fetch)
- 392
- Model card
- https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
Use cases
- Multilingual text-to-speech synthesis
- Custom voice design and cloning for specific speaker characteristics
- Audiobook and narration production pipelines
- Localization voiceover prototyping across supported languages
Pros
- Apache-2.0 license — unrestricted commercial use
- VoiceDesign capability enables custom voice profiles
- 1.7B parameter scale balances quality and inference cost
- Multilingual TTS in a single model
Cons
- 12Hz token rate is relatively low — check latency for real-time applications
- Custom voice quality depends on reference audio quality
- 1.7B model may struggle with prosody on long or complex sentences
- No pre-built speaker library — voice design requires additional tooling
Tags
qwen-ttssafetensorsqwen3_ttsaudiottsqwenmultilingualtext-to-speecharxiv:2601.15621license:apache-2.0region:us