From the model card
Fields below are copied from the tags and counters on the HuggingFace repository microsoft/Phi-4-multimodal-instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- microsoft
- Pipeline tag
- automatic-speech-recognition
- Library
- Transformers
- Weight formats
- safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- multilingual; Arabic (ar), Chinese (zh), Czech (cs), Danish (da), Dutch (nl), English (en), Finnish (fi), French (fr), German (de), Hebrew (he), Hungarian (hu), Italian (it), Japanese (ja), Korean (ko), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Spanish (es), Swedish (sv), Thai (th), Turkish (tr), Ukrainian (uk)
- Papers cited
- arXiv:2503.01743, arXiv:2407.13833
- Downloads (HF counter at last fetch)
- 462,286
- Likes (HF counter at last fetch)
- 1,612
- Model card
- https://huggingface.co/microsoft/Phi-4-multimodal-instruct
Use cases
- Multilingual multimodal assistant handling audio, image, and text
- Speech transcription and translation across 23 languages
- Visual document understanding combined with audio queries
- On-device multimodal AI in resource-constrained scenarios
Pros
- MIT license — unrestricted commercial use
- Handles audio, image, video, and text in one model
- 23-language coverage for speech and text tasks
- Phi-4-Mini base provides competitive quality for its size
Cons
- Requires custom_code — trust_remote_code=True needed
- Multi-modal routing complexity increases debugging difficulty
- Audio quality at small model scale varies on overlapping or noisy speech
- Video understanding is limited to short clip contexts
Tags
transformerssafetensorsphi4mmtext-generationnlpcodeaudioautomatic-speech-recognitionspeech-summarizationspeech-translationvisual-question-answeringphi-4-multimodalphiphi-4-minicustom_codemultilingualarzhcsda