Azure Cognitive Services
We are excited to announce Phi-4-multimodal and Phi-4-mini, the newest models in Microsoft’s Phi family of small language models. Learn more.
Model details
Phi-4-multimodal is a lightweight, 5.6-billion-parameter foundation model built to handle complex, cross-modal tasks. Designed as a versatile tool for developers, it integrates vision, audio, and text processing capabilities into a single architecture. By supporting a wide range of inputs, the model is well-suited for applications that require nuanced understanding across different media types, such as speech analysis, visual interpretation, and integrated multimodal reasoning. Its compact design allows for deployment on edge devices, making it a practical choice for IoT environments where computing power and network connectivity are constrained.
The model leverages the extensive language, vision, and speech research established in earlier Phi-3.5 and 4.0 iterations. To ensure high-quality performance and precise instruction adherence, it underwent a rigorous enhancement process that includes supervised fine-tuning, direct preference optimization, and reinforcement learning from human feedback. This lineage results in a model that balances strong reasoning and multilingual understanding with safety-conscious outputs. With its ability to process long-sequence data, it serves as a robust foundation for developers looking to build sophisticated, responsive AI agents that operate effectively across diverse linguistic and sensory domains.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Azure Cognitive Services
We are excited to announce Phi-4-multimodal and Phi-4-mini, the newest models in Microsoft’s Phi family of small language models. Learn more.
Azure Cognitive Services
Phi-4-mini brings significant enhancements in multilingual support, reasoning, and mathematics, and now, the long-awaited function calling feature is finally...
Azure Cognitive Services
The latest AI models from Phi, 4-mini-instruct and 4-multimodal-instruct, are now available in GitHub Models.
Azure Cognitive Services
Microsoft has unveiled two new additions to its Phi-4 family of small language models: Phi-4-multimodal, which integrates speech, vision, and text, and Phi-4-mini., Microsoft has unveiled two new additions to its Phi-4 family of small language models: Phi-4-multimodal, which integrates speech, vision, and text, and Phi
Azure Cognitive Services
An aggregation page on EmergentMind profiles Phi-4-multimodal-instruct as Microsoft's multimodal LLM in the Phi-4 family, implemented as a compact decoder-only architecture with modality-specific LoRA adapters and specialized encoders, projectors, and routers for efficient instruction tuning across vision-language and speech-language tasks. The page positions Phi-4-multimodal-instruct as a downstream instruction-following instantiation of the Phi-4-Multimodal design, citing the Microsoft Phi-4-Mini Technical Report and noting follow-on research (Ahn et al., September 2025) that uses the model for fine-tuning experiments including pronunciation evaluation. The summary frames Phi-4-multimodal-instruct less as an isolated checkpoint and more as an example of a broader compact multimodal design pattern.
Azure Cognitive Services
Microsoft's Phi-4-Mini Technical Report (arXiv:2503.01743v2, March 2025) introduces Phi-4-Multimodal as a compact multimodal model that integrates text, vision, and speech/audio input modalities into a single 3.8-billion-parameter checkpoint. The model uses a novel modality-extension approach built on Mixture-of-LoRAs with modality-specific routers, enabling multiple inference modes — including (vision + language), (vision + speech), and (speech/audio) — without interference between modalities. The speech/audio modality is notably lightweight, with its LoRA component containing only about 460 million parameters, yet Phi-4-Multimodal is reported as ranking first on the OpenASR leaderboard at publication. The report describes Phi-4-Mini's expanded 200K-token vocabulary for multilingual support and group query attention for efficient long-sequence generation, and outlines a separate reasoning-enhanced Phi-4-Mini preview that is not released alongside Phi-4-Mini and Phi-4-Multimodal.