From the model card
Fields below are copied from the tags and counters on the HuggingFace repository baidu/Qianfan-OCR at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- baidu
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual
- Papers cited
- arXiv:2603.13398, arXiv:2509.18189
- Downloads (HF counter at last fetch)
- 313,490
- Likes (HF counter at last fetch)
- 1,176
- Model card
- https://huggingface.co/baidu/Qianfan-OCR
Use cases
- Multilingual OCR from scanned documents and photos
- Structured information extraction from tables and forms
- Scene text recognition in natural images
- Document digitization pipeline processing diverse formats
Pros
- Apache-2.0 license
- VLM-based approach handles layout and context better than character-only OCR
- Multilingual support across major script systems
- Published model-index evaluation results for benchmarking
Cons
- qianfan_ocr model type requires specific Transformers version support
- Larger model than dedicated OCR tools — slower for simple character recognition tasks
- Custom architecture may complicate ONNX export for production
- Performance on handwriting or degraded documents not characterized
Tags
transformerssafetensorsqianfan_ocrimage-text-to-textvision-languageocrdocument-intelligenceqianfanconversationalmultilingualarxiv:2603.13398arxiv:2509.18189license:apache-2.0model-indexeval-resultsendpoints_compatibleregion:us