From the model card
Fields below are copied from the tags and counters on the HuggingFace repository openbmb/MiniCPM-o-2_6 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- openbmb
- Pipeline tag
- any-to-any
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual
- Papers cited
- arXiv:2405.17220, arXiv:2408.01800
- Datasets declared
- openbmb/RLAIF-V-Dataset
- Downloads (HF counter at last fetch)
- 398,898
- Likes (HF counter at last fetch)
- 1,296
- Model card
- https://huggingface.co/openbmb/MiniCPM-o-2_6
Use cases
- On-device voice assistants with visual context understanding
- Real-time speech-to-speech conversation on mobile hardware
- Multimodal document understanding combining OCR and NLP
- Embedded AI assistants in resource-constrained environments
Pros
- Unified speech+vision+text in a single 8B model — rare at this parameter count
- Designed for on-device deployment; quantized variants available
- Strong OCR and document understanding based on MiniCPM lineage
- Apache 2.0 license
Cons
- Audio output quality lags behind dedicated TTS systems
- 8B constrains reasoning depth on complex tasks
- Real-time streaming requires careful batching to avoid latency spikes
- Smaller community than LLaVA or Qwen-VL for troubleshooting
Tags
transformerssafetensorsminicpmofeature-extractionminicpm-oomnivisionocrmulti-imagevideocustom_codeaudiospeechvoice cloninglive Streamingrealtime speech conversationasrttsany-to-anymultilingual