From the model card
Fields below are copied from the tags and counters on the HuggingFace repository zai-org/GLM-OCR at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- zai-org
- Pipeline tag
- image-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- Chinese (zh), English (en), French (fr), Spanish (es), Russian (ru), German (de), Japanese (ja), Korean (ko)
- Papers cited
- arXiv:2603.10910
- Downloads (HF counter at last fetch)
- 2,000,195
- Likes (HF counter at last fetch)
- 2,011
- Model card
- https://huggingface.co/zai-org/GLM-OCR
Use cases
- Multilingual document text extraction from scanned PDFs
- Structured data extraction from forms and tables in images
- Receipt and invoice OCR for financial automation
- Screenshot-to-text conversion for multilingual interfaces
- Building document processing pipelines for Asian language documents
Pros
- MIT license for broad commercial use
- 8-language support including Chinese, Japanese, Korean in a single model
- Generative approach handles complex layouts better than classification-based OCR
- HuggingFace Transformers-compatible for standard inference workflows
Cons
- Generative OCR is slower than detection-based alternatives for simple text extraction
- Language coverage is limited to 8 languages — no support for Arabic, Hindi, or other scripts
- Output formatting (JSON vs. plain text) requires post-processing
- Accuracy on degraded or handwritten documents not well established
- Large model footprint vs. specialized OCR tools like Tesseract for single-language use
Tags
transformerssafetensorsglm_ocrimage-text-to-textconversationalzhenfresrudejakoarxiv:2603.10910license:miteval-resultsendpoints_compatibledeploy:azuredeploy:sagemakerregion:us