AI Tools.

Search

text to speech by Qwen

Qwen3-TTS-12Hz-1.7B-VoiceDesign

Qwen3-TTS VoiceDesign is a 1.7B text-to-speech model operating at 12Hz token rate, designed to support custom voice creation alongside standard TTS. It covers multiple languages and generates expressive speech from text input. Apache-2.0 licensed and part of Qwen's audio model family.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
text-to-speech
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual
Papers cited
arXiv:2601.15621
Downloads (HF counter at last fetch)
404,061
Likes (HF counter at last fetch)
392
Model card
https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign

Use cases

  • Multilingual text-to-speech synthesis
  • Custom voice design and cloning for specific speaker characteristics
  • Audiobook and narration production pipelines
  • Localization voiceover prototyping across supported languages

Pros

  • Apache-2.0 license — unrestricted commercial use
  • VoiceDesign capability enables custom voice profiles
  • 1.7B parameter scale balances quality and inference cost
  • Multilingual TTS in a single model

Cons

  • 12Hz token rate is relatively low — check latency for real-time applications
  • Custom voice quality depends on reference audio quality
  • 1.7B model may struggle with prosody on long or complex sentences
  • No pre-built speaker library — voice design requires additional tooling

Tags

qwen-ttssafetensorsqwen3_ttsaudiottsqwenmultilingualtext-to-speecharxiv:2601.15621license:apache-2.0region:us