From the model card
Fields below are copied from the tags and counters on the HuggingFace repository stepfun-ai/GOT-OCR2_0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- stepfun-ai
- Pipeline tag
- image-text-to-text
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual; Gothic (got)
- Papers cited
- arXiv:2409.01704, arXiv:2405.14295, arXiv:2312.06109
- Downloads (HF counter at last fetch)
- 676,295
- Likes (HF counter at last fetch)
- 1,559
- Model card
- https://huggingface.co/stepfun-ai/GOT-OCR2_0
Use cases
- End-to-end OCR on documents with mixed content — text, tables, formulas
- Mathematical formula extraction from scanned papers or textbooks
- Multilingual scene text recognition in photos or screenshots
- Building document digitization pipelines with a single model checkpoint
Pros
- 1,547 likes makes GOT-OCR2.0 one of the most validated OCR models in the community
- Handles scene text, formulas, and tables without switching models
- Multilingual support confirmed in the model card
- arxiv:2409.01704 provides detailed architecture and benchmark documentation
Cons
- Custom GOT architecture requires GOT-specific inference code
- 580M params is small — accuracy on very low-resolution or degraded input is limited
- Formula extraction accuracy depends on typesetting conventions
- Table structure understanding may miss complex merged cells or nested tables
Tags
safetensorsGOTgotvision-languageocr2.0custom_codeimage-text-to-textmultilingualarxiv:2409.01704arxiv:2405.14295arxiv:2312.06109license:apache-2.0region:us