by google
Gemma 4-E4B-IT is Google DeepMind's edge-optimized 4-billion-parameter any-to-any multimodal model from the Gemma 4 family, designed for deployment on mobile and edge devices rather than servers. The 'any-to-any' pipeline_tag indicates multimodal input and output capability beyond standard image-text-to-text. Apache 2.0 licensed.
4,740,694 ↓ · 1,530 ♡
by google
gemma-4-12B-it is Google's Gemma 4 multimodal (text + image) instruction-tuned model. It accepts both text and image inputs and produces text, making it suitable for document analysis, visual Q&A, and structured data extraction. Released under Apache-2.0, it targets users who need a capable VLM without access restrictions.
3,165,860 ↓ · 1,519 ♡
by google
This is the official Google release of Gemma 4 12B instruction-tuned in GGUF format, quantized to q4_0 using Quantization-Aware Training. Unlike community repacks, this comes directly from Google, providing clearer provenance for production pipelines that require verified model sources.
713,726 ↓ · 286 ♡
by openbmb
MiniCPM-o 2.6 is an omnimodal 8B model from OpenBMB supporting speech, image, and text inputs with real-time audio output. It targets on-device multimodal scenarios, particularly mobile and edge deployments, with end-to-end speech conversation capability.
398,898 ↓ · 1,296 ♡