From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen2.5-3B-Instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- text-generation
- Library
- Transformers
- Weight formats
- safetensors
- License tag
other— read the license file in the repo before relying on it- Lineage
-
- base model Qwen/Qwen2.5-3B
- fine-tune of Qwen/Qwen2.5-3B
- Language tags
- English (en)
- Papers cited
- arXiv:2407.10671
- Downloads (HF counter at last fetch)
- 7,557,978
- Likes (HF counter at last fetch)
- 558
- Model card
- https://huggingface.co/Qwen/Qwen2.5-3B-Instruct
Use cases
- Local inference on consumer hardware with limited VRAM
- Simple Q&A and summarization tasks where 7B is over-resourced
- API endpoint serving where latency matters more than accuracy depth
- Prototyping and development before scaling to larger models
- Batch processing simple text tasks at cost-effective throughput
Pros
- 3B scale balances quality and resource cost better than 1.5B
- Text-generation-inference compatible
- Part of maintained Qwen2.5 family
- Fits in 6-8GB VRAM at FP16 for single-consumer-GPU deployment
Cons
- License is 'other' — not Apache 2.0; verify commercial use terms
- 3B reasoning depth still limited for complex multi-step tasks
- Competitive 3B models (Phi-3.5-mini, Gemma-3-4B) should be benchmarked
- Qwen2.5 superseded by Qwen3 series — fewer ongoing optimizations
- Instruction following reliability lower than 7B+ on structured output tasks
Tags
transformerssafetensorsqwen2text-generationchatconversationalenarxiv:2407.10671base_model:Qwen/Qwen2.5-3Bbase_model:finetune:Qwen/Qwen2.5-3Blicense:othertext-generation-inferenceendpoints_compatibleregion:usdeploy:azure