AI Tools.

Search

image text to text by google

gemma-3-4b-it

Gemma 3 4B Instruct is Google's compact instruction-following model, targeting deployment on single-GPU and edge devices. It covers both text and image inputs and is suitable for conversational AI applications with moderate resource constraints.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository google/gemma-3-4b-it at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
google
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
gemma — read the license file in the repo before relying on it
Lineage
Papers cited
arXiv:1905.07830, arXiv:1905.10044, arXiv:1911.11641, arXiv:1904.09728, arXiv:1705.03551, arXiv:1911.01547, arXiv:1907.10641, arXiv:1903.00161, arXiv:2009.03300, arXiv:2304.06364, arXiv:2103.03874, arXiv:2110.14168, arXiv:2311.12022, arXiv:2108.07732, arXiv:2107.03374, arXiv:2210.03057, arXiv:2106.03193, arXiv:1910.11856, arXiv:2502.12404, arXiv:2502.21228, arXiv:2404.16816, arXiv:2104.12756, arXiv:2311.16502, arXiv:2203.10244, arXiv:2404.12390, arXiv:1810.12440, arXiv:1908.02660, arXiv:2312.11805
Downloads (HF counter at last fetch)
1,542,803
Likes (HF counter at last fetch)
1,470
Model card
https://huggingface.co/google/gemma-3-4b-it

Use cases

  • Lightweight instruction-following assistant on consumer hardware
  • Multimodal chat with image understanding at 4B scale
  • Fine-tuning base for constrained deployment scenarios
  • Embedded AI features in apps where a 7B+ model is too large

Pros

  • 4B scale runs comfortably on 8GB VRAM
  • Supports image input via Gemma 3's multimodal architecture
  • Gemma license permits commercial use with attribution
  • Google-maintained with documented benchmark results

Cons

  • Gemma license has more restrictions than Apache-2.0
  • Instruction following quality notably weaker than Llama 3.2 3B on many benchmarks
  • Vision capability is limited compared to dedicated VLMs
  • Model card has limited information on specific training data composition

Tags

transformerssafetensorsgemma3image-text-to-textconversationalarxiv:1905.07830arxiv:1905.10044arxiv:1911.11641arxiv:1904.09728arxiv:1705.03551arxiv:1911.01547arxiv:1907.10641arxiv:1903.00161arxiv:2009.03300arxiv:2304.06364arxiv:2103.03874arxiv:2110.14168arxiv:2311.12022arxiv:2108.07732arxiv:2107.03374