AI Tools.

Search

text generation by google

gemma-2b

Gemma 2B is Google's 2B-parameter open language model from early 2024, trained on 2T tokens of web, code, and math data. It was notable at release for punching above its weight class on benchmarks vs other 2B models available at the time.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository google/gemma-2b at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
google
Pipeline tag
text-generation
Library
Transformers
Weight formats
safetensors, GGUF
License tag
gemma — read the license file in the repo before relying on it
Papers cited
arXiv:2312.11805, arXiv:2009.03300, arXiv:1905.07830, arXiv:1911.11641, arXiv:1904.09728, arXiv:1905.10044, arXiv:1907.10641, arXiv:1811.00937, arXiv:1809.02789, arXiv:1911.01547, arXiv:1705.03551, arXiv:2107.03374, arXiv:2108.07732, arXiv:2110.14168, arXiv:2304.06364, arXiv:2206.04615, arXiv:1804.06876, arXiv:2110.08193, arXiv:2009.11462, arXiv:2101.11718, arXiv:1804.09301, arXiv:2109.07958, arXiv:2203.09509
Downloads (HF counter at last fetch)
307,263
Likes (HF counter at last fetch)
1,185
Model card
https://huggingface.co/google/gemma-2b

Use cases

  • On-device LLM inference where 7B is too large
  • Text generation and chat with a small open model
  • Fine-tuning base for custom domain adaptation at 2B scale
  • Educational use for learning model fine-tuning workflows

Pros

  • Competitive benchmark scores for 2B class at release
  • Apache 2.0-compatible Gemma license for research and commercial use (with conditions)
  • Google quality control in pretraining pipeline
  • Strong starting point for low-resource fine-tuning

Cons

  • Gemma 2B is now outperformed by Gemma 2B-it, Phi-3-mini, and Qwen2-1.5B on most tasks
  • Base model (not instruct) requires prompt engineering or fine-tuning for chat use
  • 2B knowledge depth is shallow — hallucination rate is higher than 7B+ models
  • Knowledge cutoff early 2024

Tags

transformerssafetensorsggufgemmatext-generationarxiv:2312.11805arxiv:2009.03300arxiv:1905.07830arxiv:1911.11641arxiv:1904.09728arxiv:1905.10044arxiv:1907.10641arxiv:1811.00937arxiv:1809.02789arxiv:1911.01547arxiv:1705.03551arxiv:2107.03374arxiv:2108.07732arxiv:2110.14168arxiv:2304.06364