AI Tools.

Search

automatic speech recognition by microsoft

Phi-4-multimodal-instruct

Phi-4-Multimodal-Instruct is Microsoft's compact multimodal model handling text, audio, images, and video in a single instruction-tuned model. Based on Phi-4-Mini, it covers 23 languages and supports speech recognition, speech translation, and visual QA. MIT-licensed — fully permissive for commercial use.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository microsoft/Phi-4-multimodal-instruct at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
microsoft
Pipeline tag
automatic-speech-recognition
Library
Transformers
Weight formats
safetensors
License tag
mit — read the license file in the repo before relying on it
Language tags
multilingual; Arabic (ar), Chinese (zh), Czech (cs), Danish (da), Dutch (nl), English (en), Finnish (fi), French (fr), German (de), Hebrew (he), Hungarian (hu), Italian (it), Japanese (ja), Korean (ko), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Spanish (es), Swedish (sv), Thai (th), Turkish (tr), Ukrainian (uk)
Papers cited
arXiv:2503.01743, arXiv:2407.13833
Downloads (HF counter at last fetch)
462,286
Likes (HF counter at last fetch)
1,612
Model card
https://huggingface.co/microsoft/Phi-4-multimodal-instruct

Use cases

  • Multilingual multimodal assistant handling audio, image, and text
  • Speech transcription and translation across 23 languages
  • Visual document understanding combined with audio queries
  • On-device multimodal AI in resource-constrained scenarios

Pros

  • MIT license — unrestricted commercial use
  • Handles audio, image, video, and text in one model
  • 23-language coverage for speech and text tasks
  • Phi-4-Mini base provides competitive quality for its size

Cons

  • Requires custom_code — trust_remote_code=True needed
  • Multi-modal routing complexity increases debugging difficulty
  • Audio quality at small model scale varies on overlapping or noisy speech
  • Video understanding is limited to short clip contexts

Tags

transformerssafetensorsphi4mmtext-generationnlpcodeaudioautomatic-speech-recognitionspeech-summarizationspeech-translationvisual-question-answeringphi-4-multimodalphiphi-4-minicustom_codemultilingualarzhcsda