From the model card
Fields below are copied from the tags and counters on the HuggingFace repository openbmb/MiniCPM-V-4_5 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- openbmb
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual
- Papers cited
- arXiv:2509.18154, arXiv:2403.11703
- Datasets declared
- openbmb/RLAIF-V-Dataset
- Downloads (HF counter at last fetch)
- 389,181
- Likes (HF counter at last fetch)
- 1,098
- Model card
- https://huggingface.co/openbmb/MiniCPM-V-4_5
Use cases
- Efficient visual question answering on consumer hardware
- OCR and document understanding from image inputs
- Multi-image comparison tasks in research workflows
- Video understanding at low inference cost
Pros
- 1,094 likes confirm strong community validation of quality vs. cost trade-off
- Multi-image and video inputs expand coverage beyond single-frame VL models
- Optimized for low VRAM deployment relative to its capability level
- OpenBMB publishes detailed evaluation across OCR, VQA, and chart tasks
Cons
- Custom minicpm-v architecture requires custom inference code
- Video processing throughput is slower than dedicated video understanding models
- Accuracy on high-resolution images degrades compared to larger VL models
- Custom code dependency complicates integration into generic VL serving stacks
Tags
transformerssafetensorsminicpmvfeature-extractionminicpm-vvisionocrmulti-imagevideocustom_codeimage-text-to-textconversationalmultilingualdataset:openbmb/RLAIF-V-Datasetarxiv:2509.18154arxiv:2403.11703license:apache-2.0eval-resultsregion:us