Single model supporting speech, vision and text understanding.
- Feb 26, 2025
- —
- textimageaudio
- —
- —
- 1
Releases
Feb 26, 2025Phi-4 Multimodal releasedOpen weights
Microsoft releases Phi-4 Multimodal, a 5.6-billion-parameter model jointly processing speech, vision and text.

