From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen2-VL-7B-Instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model Qwen/Qwen2-VL-7B
- fine-tune of Qwen/Qwen2-VL-7B
- Language tags
- English (en)
- Papers cited
- arXiv:2409.12191, arXiv:2308.12966
- Downloads (HF counter at last fetch)
- 1,254,568
- Likes (HF counter at last fetch)
- 1,285
- Model card
- https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct
Use cases
- Document understanding and OCR from scanned images
- Visual question answering over charts and figures
- Screenshot-to-code or UI description tasks
- Multi-image reasoning in a single context window
Pros
- Native variable-resolution input without cropping
- Strong OCR and document parsing compared to peers
- Apache-2.0 license permits commercial use
- Active inference support on vLLM and text-generation-inference
Cons
- 7B scale still struggles with fine-grained spatial reasoning
- Hallucination rate higher than GPT-4V on knowledge-grounded tasks
- Requires ~16GB VRAM for full bfloat16 serving
- No audio modality despite the VL name
Tags
transformerssafetensorsqwen2_vlimage-text-to-textmultimodalconversationalenarxiv:2409.12191arxiv:2308.12966base_model:Qwen/Qwen2-VL-7Bbase_model:finetune:Qwen/Qwen2-VL-7Blicense:apache-2.0eval-resultstext-generation-inferenceendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure