AI Tools.

Search

text generation by google

gemma-2-2b-it

Gemma 2 2B Instruct is Google's smallest instruction-tuned model in the Gemma 2 family, using the same sliding window + full attention hybrid and logit soft-capping as the 9B variant but at 2.6 billion parameters. At release it set a new bar for sub-3B instruction models on standard benchmarks. It is Apache 2.0 licensed and runs on consumer hardware.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository google/gemma-2-2b-it at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
google
Pipeline tag
text-generation
Library
Transformers
Weight formats
safetensors
License tag
gemma — read the license file in the repo before relying on it
Lineage
Papers cited
arXiv:2009.03300, arXiv:1905.07830, arXiv:1911.11641, arXiv:1904.09728, arXiv:1905.10044, arXiv:1907.10641, arXiv:1811.00937, arXiv:1809.02789, arXiv:1911.01547, arXiv:1705.03551, arXiv:2107.03374, arXiv:2108.07732, arXiv:2110.14168, arXiv:2009.11462, arXiv:2101.11718, arXiv:2110.08193, arXiv:1804.09301, arXiv:2109.07958, arXiv:1804.06876, arXiv:2103.03874, arXiv:2304.06364, arXiv:1903.00161, arXiv:2206.04615, arXiv:2203.09509, arXiv:2403.13793
Downloads (HF counter at last fetch)
676,599
Likes (HF counter at last fetch)
1,480
Model card
https://huggingface.co/google/gemma-2-2b-it

Use cases

  • On-device or embedded assistant where the 9B model is too large
  • Lightweight summarisation and question answering in production
  • Fine-tuning baseline for narrow-domain instruction following at minimal cost
  • Running instruction-following on hardware with 4-6GB VRAM
  • Batch offline processing where latency matters more than peak quality

Pros

  • Best sub-3B instruction model at time of release; still competitive in its class
  • Apache 2.0 license; TGI and Azure deployment supported
  • 1359 likes; the most widely adopted Gemma 2B variant
  • Sliding window attention improves coherence on longer contexts at small scale

Cons

  • 2B capacity makes it unsuitable for complex reasoning or factual queries
  • Gemma 2 2B is now outpaced by SmolLM2 and Qwen3-0.6/1.5B in size/quality trade-off
  • Soft-capping can produce overconfident outputs on borderline knowledge questions
  • Limited multilingual capability despite some cross-lingual fine-tuning

Tags

transformerssafetensorsgemma2text-generationconversationalarxiv:2009.03300arxiv:1905.07830arxiv:1911.11641arxiv:1904.09728arxiv:1905.10044arxiv:1907.10641arxiv:1811.00937arxiv:1809.02789arxiv:1911.01547arxiv:1705.03551arxiv:2107.03374arxiv:2108.07732arxiv:2110.14168arxiv:2009.11462arxiv:2101.11718