AI Tools.

Search

automatic speech recognition by Qwen

Qwen3-ForcedAligner-0.6B

Qwen3-ForcedAligner-0.6B is a forced alignment model from the Qwen3 ASR family, designed to align audio segments to text transcripts at the phoneme or word level. At 0.6B parameters it's compact for deployment in audio processing pipelines. Apache-2.0 licensed.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3-ForcedAligner-0.6B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
automatic-speech-recognition
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Papers cited
arXiv:2601.21337
Downloads (HF counter at last fetch)
589,084
Likes (HF counter at last fetch)
156
Model card
https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B

Use cases

  • Word-level timestamp generation from audio and transcript pairs
  • Subtitle synchronization from transcripts
  • Training data alignment for TTS and ASR model training
  • Forced phoneme alignment for pronunciation assessment

Pros

  • Apache-2.0 license
  • Compact 0.6B size for a forced alignment task
  • Integrates with Qwen3 audio ecosystem
  • safetensors format

Cons

  • 0.6B may struggle with rapid speech or heavily accented audio
  • No published alignment accuracy benchmarks on standard datasets
  • Limited to supported Qwen3 ASR languages
  • Forced alignment is a narrow task — not suitable for general ASR transcription

Tags

safetensorsqwen3_asrautomatic-speech-recognitionarxiv:2601.21337license:apache-2.0region:us