From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen2-Audio-7B-Instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- audio-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:2407.10759, arXiv:2311.07919
- Downloads (HF counter at last fetch)
- 377,338
- Likes (HF counter at last fetch)
- 551
- Model card
- https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct
Use cases
- Audio understanding and description from arbitrary audio clips
- Speech transcription paired with contextual Q&A
- Meeting summarization from audio input
- Multilingual speech analysis combining audio and text instructions
Pros
- Apache-2.0 license
- Handles both audio understanding and speech recognition in one model
- Instruction-tuned for conversational audio Q&A
- Transformers-compatible with standard qwen2_audio pipeline
Cons
- 7B parameters means substantial VRAM requirement for audio+text inference
- Audio processing adds latency compared to text-only models
- Performance on music, environmental sounds, or non-speech audio varies significantly
- Limited fine-tuning recipes for audio domain adaptation
Tags
transformerssafetensorsqwen2_audiotext2text-generationchataudioaudio-text-to-textenarxiv:2407.10759arxiv:2311.07919license:apache-2.0endpoints_compatibleregion:us