From the model card
Fields below are copied from the tags and counters on the HuggingFace repository meta-llama/Llama-3.2-3B-Instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- meta-llama
- Pipeline tag
- text-generation
- Library
- Transformers
- Framework tags
- PyTorch
- Weight formats
- safetensors
- License tag
llama3.2— read the license file in the repo before relying on it- Language tags
- English (en), German (de), French (fr), Italian (it), Portuguese (pt), Hindi (hi), Spanish (es), Thai (th)
- Papers cited
- arXiv:2204.05149, arXiv:2405.16406
- Downloads (HF counter at last fetch)
- 1,385,210
- Likes (HF counter at last fetch)
- 2,504
- Model card
- https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct
Use cases
- On-device assistant on phones and laptops
- Lightweight server-side inference where cost-per-token matters most
- Instruction following in latency-sensitive APIs
- Base model for fine-tuning narrow instruction tasks
Pros
- 3B scale offers good quality-per-FLOP for its class
- Llama license permits broad commercial use
- Strong on instruction following benchmarks relative to size
- Excellent llama.cpp and Ollama support with multiple quantizations
Cons
- 3B capacity means weaker multi-step reasoning than 7B+ models
- Knowledge depth limited compared to larger Llama variants
- Context length shorter than Llama 3.1 8B
- Hallucination rate noticeably higher than 7–8B models on factual queries
Tags
transformerssafetensorsllamatext-generationfacebookmetapytorchllama-3conversationalendefritpthiestharxiv:2204.05149arxiv:2405.16406license:llama3.2