From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3-ASR-1.7B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- automatic-speech-recognition
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Papers cited
- arXiv:2601.21337
- Downloads (HF counter at last fetch)
- 4,206,898
- Likes (HF counter at last fetch)
- 1,066
- Model card
- https://huggingface.co/Qwen/Qwen3-ASR-1.7B
Use cases
- Multilingual speech transcription for meeting and podcast tools
- ASR in production pipelines where Whisper-large is too slow
- Chinese-English bilingual transcription
- Building voice interfaces for Qwen-based LLM applications
Pros
- 1.7B scale balances ASR quality and inference cost
- Multilingual with emphasis on Chinese and English
- Apache-2.0 licensed
- Designed for integration with Qwen LLM ecosystem
Cons
- Benchmark WER comparisons against Whisper-large-v3 not yet widely published
- 1.7B is heavier than Whisper-small for similar or lower quality on English-only tasks
- Streaming inference support documentation sparse
- Less tested on accented speech and low-resource languages than Whisper
Tags
safetensorsqwen3_asrautomatic-speech-recognitionarxiv:2601.21337license:apache-2.0eval-resultsdeploy:sagemakerdeploy:azureregion:us