AI Tools.

Search

image to text models

2 models · ranked by HuggingFace downloads

GLM-OCR

by zai-org

GLM-OCR is a multilingual OCR and document understanding model from ZhipuAI, built on the GLM architecture and supporting text recognition across Chinese, English, French, Spanish, Russian, German, Japanese, and Korean. It treats OCR as a sequence generation task, enabling structured text extraction from document images and screenshots. MIT licensed.

2,000,195 ↓ · 2,011 ♡

blip-image-captioning-base

by Salesforce

BLIP (Bootstrapped Language-Image Pretraining) base model for image captioning, using a vision encoder connected to a decoder via cross-attention. It introduced a bootstrapping approach that filters noisy web-crawled image-text pairs during training.

1,823,300 ↓ · 886 ♡