AI Tools.

Search

image text to text by baidu

Qianfan-OCR

Qianfan-OCR is Baidu's vision-language model specialized for optical character recognition and document intelligence, supporting multilingual text extraction from images. It combines a vision encoder with a language model for scene text understanding beyond simple character recognition. Apache-2.0 licensed with published benchmark results.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository baidu/Qianfan-OCR at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
baidu
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Language tags
multilingual
Papers cited
arXiv:2603.13398, arXiv:2509.18189
Downloads (HF counter at last fetch)
313,490
Likes (HF counter at last fetch)
1,176
Model card
https://huggingface.co/baidu/Qianfan-OCR

Use cases

  • Multilingual OCR from scanned documents and photos
  • Structured information extraction from tables and forms
  • Scene text recognition in natural images
  • Document digitization pipeline processing diverse formats

Pros

  • Apache-2.0 license
  • VLM-based approach handles layout and context better than character-only OCR
  • Multilingual support across major script systems
  • Published model-index evaluation results for benchmarking

Cons

  • qianfan_ocr model type requires specific Transformers version support
  • Larger model than dedicated OCR tools — slower for simple character recognition tasks
  • Custom architecture may complicate ONNX export for production
  • Performance on handwriting or degraded documents not characterized

Tags

transformerssafetensorsqianfan_ocrimage-text-to-textvision-languageocrdocument-intelligenceqianfanconversationalmultilingualarxiv:2603.13398arxiv:2509.18189license:apache-2.0model-indexeval-resultsendpoints_compatibleregion:us