From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3-ForcedAligner-0.6B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- automatic-speech-recognition
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Papers cited
- arXiv:2601.21337
- Downloads (HF counter at last fetch)
- 589,084
- Likes (HF counter at last fetch)
- 156
- Model card
- https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B
Use cases
- Word-level timestamp generation from audio and transcript pairs
- Subtitle synchronization from transcripts
- Training data alignment for TTS and ASR model training
- Forced phoneme alignment for pronunciation assessment
Pros
- Apache-2.0 license
- Compact 0.6B size for a forced alignment task
- Integrates with Qwen3 audio ecosystem
- safetensors format
Cons
- 0.6B may struggle with rapid speech or heavily accented audio
- No published alignment accuracy benchmarks on standard datasets
- Limited to supported Qwen3 ASR languages
- Forced alignment is a narrow task — not suitable for general ASR transcription
Tags
safetensorsqwen3_asrautomatic-speech-recognitionarxiv:2601.21337license:apache-2.0region:us