AI Tools.

Search

image text to text by Qwen

Qwen2-VL-2B-Instruct

Qwen2-VL-2B-Instruct is a 2B parameter vision-language model from Alibaba's Qwen team, supporting image and video understanding alongside text instruction-following. At 2B parameters it runs on consumer GPUs while retaining competitive OCR, chart reading, and visual QA accuracy. It is the instruction-tuned version of the Qwen2-VL-2B base.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen2-VL-2B-Instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Language tags
English (en)
Papers cited
arXiv:2409.12191, arXiv:2308.12966
Downloads (HF counter at last fetch)
1,821,310
Likes (HF counter at last fetch)
518
Model card
https://huggingface.co/Qwen/Qwen2-VL-2B-Instruct

Use cases

  • Captioning product images in e-commerce pipelines
  • Visual question answering over uploaded charts or diagrams
  • Document OCR on edge devices with limited VRAM
  • Lightweight VQA in mobile or embedded applications

Pros

  • Runs in under 8GB VRAM making it edge-deployable
  • Apache 2.0 license with no commercial restrictions
  • Strong OCR and structured document understanding for its parameter count

Cons

  • 2B scale trails larger VL models on complex visual reasoning tasks
  • Shorter context window than Qwen2-VL-7B variant
  • Video understanding limited compared to dedicated video-language models

Tags

transformerssafetensorsqwen2_vlimage-text-to-textmultimodalconversationalenarxiv:2409.12191arxiv:2308.12966base_model:Qwen/Qwen2-VL-2Bbase_model:finetune:Qwen/Qwen2-VL-2Blicense:apache-2.0text-generation-inferenceendpoints_compatibleregion:usdeploy:azure