AI Tools.

Search

image to text by zai-org

GLM-OCR

GLM-OCR is a multilingual OCR and document understanding model from ZhipuAI, built on the GLM architecture and supporting text recognition across Chinese, English, French, Spanish, Russian, German, Japanese, and Korean. It treats OCR as a sequence generation task, enabling structured text extraction from document images and screenshots. MIT licensed.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository zai-org/GLM-OCR at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
zai-org
Pipeline tag
image-to-text
Library
Transformers
Weight formats
safetensors
License tag
mit — read the license file in the repo before relying on it
Language tags
Chinese (zh), English (en), French (fr), Spanish (es), Russian (ru), German (de), Japanese (ja), Korean (ko)
Papers cited
arXiv:2603.10910
Downloads (HF counter at last fetch)
2,000,195
Likes (HF counter at last fetch)
2,011
Model card
https://huggingface.co/zai-org/GLM-OCR

Use cases

  • Multilingual document text extraction from scanned PDFs
  • Structured data extraction from forms and tables in images
  • Receipt and invoice OCR for financial automation
  • Screenshot-to-text conversion for multilingual interfaces
  • Building document processing pipelines for Asian language documents

Pros

  • MIT license for broad commercial use
  • 8-language support including Chinese, Japanese, Korean in a single model
  • Generative approach handles complex layouts better than classification-based OCR
  • HuggingFace Transformers-compatible for standard inference workflows

Cons

  • Generative OCR is slower than detection-based alternatives for simple text extraction
  • Language coverage is limited to 8 languages — no support for Arabic, Hindi, or other scripts
  • Output formatting (JSON vs. plain text) requires post-processing
  • Accuracy on degraded or handwritten documents not well established
  • Large model footprint vs. specialized OCR tools like Tesseract for single-language use

Tags

transformerssafetensorsglm_ocrimage-text-to-textconversationalzhenfresrudejakoarxiv:2603.10910license:miteval-resultsendpoints_compatibledeploy:azuredeploy:sagemakerregion:us