AI Tools.

Search

image text to text models

36 models · ranked by HuggingFace downloads

Qwen3.6-35B-A3B-FP8

by Qwen

FP8-quantized version of Qwen3.6-35B-A3B for deployment on hardware with FP8 support (H100/H200). Reduces memory footprint and inference latency compared to BF16 with minimal quality degradation on most benchmarks.

13,251,463 ↓ · 372 ♡

Qwen3.5-9B

by Qwen

Qwen3.5-9B is a 9-billion-parameter instruction-tuned vision-language model from Alibaba Cloud's Qwen3.5 series, fine-tuned from Qwen3.5-9B-Base for multimodal conversational tasks. It accepts image and text inputs for visual reasoning, document understanding, and grounded question answering. Apache 2.0 licensed.

12,184,741 ↓ · 1,895 ♡

Qwen3-VL-8B-Instruct

by Qwen

Qwen3-VL-8B-Instruct is Alibaba Cloud's 8-billion-parameter vision-language model from the Qwen3-VL series, extending the VL line with improved visual reasoning and document understanding. It targets mid-tier server GPU deployment where 2B VLMs are insufficient and 30B+ is impractical. Apache 2.0 licensed.

9,810,595 ↓ · 1,081 ♡

gemma-4-31B-it

by google

Gemma 4-31B-IT is Google DeepMind's 31-billion-parameter instruction-tuned vision-language model from the Gemma 4 family, supporting both image and text inputs. It offers strong multimodal reasoning at open-weight scale, with Apache 2.0 licensing making it directly deployable for commercial applications. Part of the gemma4 architecture with improvements over Gemma 2.

7,973,406 ↓ · 3,712 ♡

gemma-4-26B-A4B-it

by google

Gemma 4-26B-A4B-IT is Google DeepMind's 26-billion-total-parameter MoE (Mixture-of-Experts) vision-language model, with approximately 4 billion active parameters per token. The MoE design means it achieves 26B parameter quality while activating only ~4B per forward pass, reducing per-token compute relative to a dense 26B model. Apache 2.0 licensed.

7,837,690 ↓ · 1,476 ♡

Qwen2.5-VL-7B-Instruct

by Qwen

Qwen2.5-VL-7B-Instruct is Alibaba Cloud's 7-billion-parameter vision-language model from the Qwen2.5-VL series, accepting image and video inputs alongside text for visual question answering, document understanding, and grounding tasks. It supports multiple image resolutions dynamically and shows improved OCR and document reasoning compared to the earlier Qwen-VL series. Apache 2.0 licensed.

7,720,403 ↓ · 1,697 ♡

Qwen3.6-27B-FP8

by Qwen

FP8-quantized version of Qwen 3.6 27B for H100/H200 serving. Reduces memory from ~54GB (BF16) to approximately 27GB while maintaining near-BF16 quality on most benchmarks for a dense multimodal model.

7,537,356 ↓ · 353 ♡

Qwen3.8-27B

by Qwen

Alibaba's 27B multimodal model that accepts images alongside text prompts in a single BF16 checkpoint. Competitive with similarly sized models on vision benchmarks and code tasks. Apache 2.0 licensed with Azure deploy integration.

5,254,882 ↓ · 13,850 ♡

Qwen3.6-35B-A3B

by Qwen

Qwen 3.6 is a Mixture-of-Experts model with 35B total parameters but only 3B active per token, giving MoE inference efficiency at near-35B capacity. It handles image and text inputs and is competitive with dense 14–20B models on standard benchmarks.

4,565,289 ↓ · 2,771 ♡

Unlimited-OCR

by baidu

Baidu's Unlimited-OCR is a vision-language model targeting text recognition across multiple scripts, layouts, and document types. Accompanies a preprint (arXiv:2606.23050) and ships with published eval results on standard OCR benchmarks.

3,008,635 ↓ · 4,182 ♡

Qwen3-VL-2B-Instruct

by Qwen

Qwen3-VL-2B-Instruct is a 2-billion-parameter vision-language model from Alibaba Cloud that jointly processes images and text for visual question answering, captioning, and document understanding. Its 2B scale positions it as one of the smaller instruction-tuned VLMs capable of zero-shot visual reasoning. Apache 2.0 licensed.

2,876,532 ↓ · 457 ♡

Kimi-K3

by moonshotai

Kimi-K3 is Moonshot AI's large-scale multimodal model designed for image-text understanding and reasoning tasks. With over 9,000 community likes it is among the most widely adopted recent open-weight multimodal releases. Covers both visual comprehension and language reasoning in a single model. Non-standard license — check Moonshot AI's terms.

2,617,373 ↓ · 11,177 ♡

DeepSeek-OCR

by deepseek-ai

DeepSeek OCR is a vision-language model from DeepSeek optimized specifically for optical character recognition from natural scene and document images. It aims to handle mixed layouts, multi-language text, and complex typographic scenarios.

2,382,527 ↓ · 3,351 ♡

Qwen3.5-27B

by Qwen

Qwen 3.5 27B is a dense image-text-to-text model from Alibaba, positioned between the 14B and 72B variants for users who need more capacity than 14B but can't serve 72B. It handles both vision and language instructions.

2,369,643 ↓ · 1,040 ♡

Qwen3.5-35B-A3B

by Qwen

Qwen3.5-35B-A3B is a 35B total parameter mixture-of-experts multimodal model from Alibaba, with approximately 3B active parameters per token during inference. It combines vision and language understanding for image captioning, visual QA, and document analysis tasks at lower compute cost than a dense 35B model. Apache 2.0 licensed.

2,362,260 ↓ · 1,498 ♡

llava-1.5-7b-hf

by llava-hf

LLaVA 1.5 7B connects a CLIP ViT-L/14@336 vision encoder to Vicuna 7B via a simple MLP projection. It was a state-of-the-art open multimodal model at release and remains widely used as a baseline for vision-language research.

1,927,882 ↓ · 372 ♡

Qwen2-VL-2B-Instruct

by Qwen

Qwen2-VL-2B-Instruct is a 2B parameter vision-language model from Alibaba's Qwen team, supporting image and video understanding alongside text instruction-following. At 2B parameters it runs on consumer GPUs while retaining competitive OCR, chart reading, and visual QA accuracy. It is the instruction-tuned version of the Qwen2-VL-2B base.

1,821,310 ↓ · 518 ♡

moondream2

by vikhyatk

Moondream2 is a 1.9B parameter vision-language model designed to be the smallest model that can meaningfully answer questions about images. It pairs a SigLIP vision encoder with a Phi-1.5 language backbone and achieves surprising capability at its size.

1,605,038 ↓ · 1,435 ♡

gemma-3-4b-it

by google

Gemma 3 4B Instruct is Google's compact instruction-following model, targeting deployment on single-GPU and edge devices. It covers both text and image inputs and is suitable for conversational AI applications with moderate resource constraints.

1,542,803 ↓ · 1,470 ♡

Qwen2-VL-7B-Instruct

by Qwen

Qwen2-VL 7B is Alibaba's second-generation vision-language model, instruction-tuned to follow text+image prompts. It handles variable-resolution inputs natively and scores competitively against GPT-4V on standard multimodal benchmarks at the 7B scale.

1,254,568 ↓ · 1,285 ♡

diffusiongemma-26B-A4B-it

by google

diffusiongemma-26B-A4B-it is Google's experimental diffusion-based language model built on the Gemma 4 MoE architecture, applying masked diffusion to text generation instead of autoregressive decoding. At 26B active-parameter scale it explores whether diffusion LMs can match autoregressive quality on instruction-following tasks. It accepts text and image inputs and produces text through iterative denoising.

1,170,594 ↓ · 1,204 ♡

medgemma-4b-it

by google

MedGemma-4B-it is Google's 4B instruction-tuned multimodal model specialized for medical image and text understanding, covering radiology, dermatology, pathology, and ophthalmology. It accepts medical images (chest X-rays, skin images, histology slides, fundus photos) paired with clinical questions. Not cleared for clinical decision support — research and development only.

1,012,002 ↓ · 1,043 ♡

Qwen3.5-4B-AWQ-4bit

by cyankiwi

AWQ 4-bit quantization of Qwen3.5-4B, a dense multimodal model supporting image-text-to-text tasks. At 4B parameters with AWQ compression, inference fits within ~4 GB VRAM, making it accessible on mid-range consumer cards. compressed-tensors format targets vLLM serving.

713,704 ↓ · 21 ♡

GOT-OCR2_0

by stepfun-ai

GOT-OCR2.0 (General OCR Theory) is a 580M-parameter image-text-to-text model from UCAS that unifies diverse OCR tasks under a single architecture. With 1,547 likes it is among the most popular specialized OCR models on HuggingFace, supporting formula, table, and scene text recognition.

676,295 ↓ · 1,559 ♡

Muse-Glimmer-30B

by meta-models

Muse-Glimmer-30B is a 30B multimodal model from meta-models (unaffiliated with Meta AI) combining vision and text understanding. With two published arXiv papers and Azure deploy integration, it targets researchers wanting an Apache-2.0-licensed alternative to commercially-gated frontier VLMs.

623,035 ↓ · 1,845 ♡

MiniCPM-V-4.6

by openbmb

MiniCPM-V-4.6 is OpenBMB's MiniCPM-V 4.6, a lightweight on-device multimodal model optimized for image+text tasks at minimal parameter count. Version 4.6 targets improved document OCR, mathematical diagram understanding, and multilingual captioning within the constraints of mobile or edge deployment. It is compatible with deployment via llama.cpp or the MiniCPM-specific inference stack.

523,553 ↓ · 1,206 ♡

Kimi-K2.7-Code

by moonshotai

Kimi-K2.7-Code is Moonshot AI's code-focused multimodal model built on the kimi_k25 architecture, accepting both image and text inputs. It uses compressed-tensors for efficient weight storage and exposes custom model code, indicating non-standard architectural components beyond base Transformers.

440,001 ↓ · 1,370 ♡

Llama-4-Scout-17B-16E-Instruct

by meta-llama

Llama 4 Scout is Meta's first MoE entry in the Llama series: 17B parameters per expert across 16 experts, with a small number active per token. The instruct variant follows instructions and handles image-text inputs natively, supporting 12 languages. Scout targets deployments where multimodal capability is needed at a lower active-parameter cost than dense Llama 3 models.

411,931 ↓ · 1,333 ♡

granite-docling-258M

by ibm-granite

granite-docling-258M is a 258M-parameter vision-language model fine-tuned specifically for document understanding tasks within the Docling pipeline. It handles OCR, layout parsing, table extraction, formula recognition, and chart reading in a single inference pass. The model is built on the Idefics3 architecture and integrates directly with the open-source Docling library.

407,501 ↓ · 1,254 ♡

dots.ocr

by rednote-hilab

dots.ocr is RedNote's specialized OCR model for structured document parsing, capable of extracting text from complex layouts including tables, mathematical formulas, and mixed Chinese-English documents. The 1315 community likes reflect substantial real-world adoption for document digitization use cases.

394,815 ↓ · 1,315 ♡

MiniCPM-V-4_5

by openbmb

OpenBMB's MiniCPM-V-4.5 is an efficient multimodal vision-language model supporting single and multi-image inputs as well as video. With 1,094 likes it is one of the most community-validated efficient VL models, excelling in OCR, chart understanding, and visual question answering.

389,181 ↓ · 1,098 ♡

LocateAnything-3B

by nvidia

LocateAnything-3B is a 3-billion-parameter vision-language model from NVIDIA that performs open-vocabulary object grounding and localization via natural language queries. It is fine-tuned from Qwen2.5-3B-Instruct using NVIDIA's Eagle visual encoder framework and targets conversational grounding workflows. The model supports referring expression comprehension and visual question answering with spatial outputs.

384,700 ↓ · 2,904 ♡

Qwen3.5-9B-AWQ-4bit

by cyankiwi

AWQ 4-bit quantization of Qwen3.5-9B, a dense image-text-to-text model. At 9B parameters with AWQ INT4, inference requires roughly 6-8 GB VRAM, placing it within reach of RTX 3080/4070-class cards. compressed-tensors format is vLLM-native.

363,295 ↓ · 36 ♡

dots.ocr

by dots-studio

dots.ocr is an image-text-to-text model specializing in optical character recognition with layout understanding, table extraction, and mathematical formula parsing. With 1,318 likes and 452K downloads it has strong community adoption for structured document digitization.

358,798 ↓ · 1,321 ♡

Qwen3.5-27B-GPTQ-Int4

by Qwen

Official Alibaba GPTQ INT4 quantization of Qwen3.5-27B, a dense multimodal model for image and text tasks. GPTQ INT4 reduces memory to approximately 15-18 GB, making the model accessible on A100 or RTX 4090-class hardware. Apache-2.0 licensed.

354,659 ↓ · 55 ♡

Qianfan-OCR

by baidu

Qianfan-OCR is Baidu's vision-language model specialized for optical character recognition and document intelligence, supporting multilingual text extraction from images. It combines a vision encoder with a language model for scene text understanding beyond simple character recognition. Apache-2.0 licensed with published benchmark results.

313,490 ↓ · 1,176 ♡