AI Tools.

Search

image text to text by stepfun-ai

GOT-OCR2_0

GOT-OCR2.0 (General OCR Theory) is a 580M-parameter image-text-to-text model from UCAS that unifies diverse OCR tasks under a single architecture. With 1,547 likes it is among the most popular specialized OCR models on HuggingFace, supporting formula, table, and scene text recognition.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository stepfun-ai/GOT-OCR2_0 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
stepfun-ai
Pipeline tag
image-text-to-text
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual; Gothic (got)
Papers cited
arXiv:2409.01704, arXiv:2405.14295, arXiv:2312.06109
Downloads (HF counter at last fetch)
676,295
Likes (HF counter at last fetch)
1,559
Model card
https://huggingface.co/stepfun-ai/GOT-OCR2_0

Use cases

  • End-to-end OCR on documents with mixed content — text, tables, formulas
  • Mathematical formula extraction from scanned papers or textbooks
  • Multilingual scene text recognition in photos or screenshots
  • Building document digitization pipelines with a single model checkpoint

Pros

  • 1,547 likes makes GOT-OCR2.0 one of the most validated OCR models in the community
  • Handles scene text, formulas, and tables without switching models
  • Multilingual support confirmed in the model card
  • arxiv:2409.01704 provides detailed architecture and benchmark documentation

Cons

  • Custom GOT architecture requires GOT-specific inference code
  • 580M params is small — accuracy on very low-resolution or degraded input is limited
  • Formula extraction accuracy depends on typesetting conventions
  • Table structure understanding may miss complex merged cells or nested tables

Tags

safetensorsGOTgotvision-languageocr2.0custom_codeimage-text-to-textmultilingualarxiv:2409.01704arxiv:2405.14295arxiv:2312.06109license:apache-2.0region:us