Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Azure Cognitive Services logo

Model details

Phi-4-multimodal

Phi-4-multimodal is a lightweight, 5.6-billion-parameter foundation model built to handle complex, cross-modal tasks. Designed as a versatile tool for developers, it integrates vision, audio, and text processing capabilities into a single architecture. By supporting a wide range of inputs, the model is well-suited for applications that require nuanced understanding across different media types, such as speech analysis, visual interpretation, and integrated multimodal reasoning. Its compact design allows for deployment on edge devices, making it a practical choice for IoT environments where computing power and network connectivity are constrained.

The model leverages the extensive language, vision, and speech research established in earlier Phi-3.5 and 4.0 iterations. To ensure high-quality performance and precise instruction adherence, it underwent a rigorous enhancement process that includes supervised fine-tuning, direct preference optimization, and reinforcement learning from human feedback. This lineage results in a model that balances strong reasoning and multilingual understanding with safety-conscious outputs. With its ability to process long-sequence data, it serves as a robust foundation for developers looking to build sophisticated, responsive AI agents that operate effectively across diverse linguistic and sensory domains.

Azure Cognitive Servicesphi-4-multimodalphi

Quick Info

Powered by
Provider
Azure Cognitive Services
Model key
phi-4-multimodal
Release date
Dec 11, 2024
Last updated
Dec 11, 2024
Knowledge cutoff
2023-10
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.08
Output token cost
$0.32

Limits

Output tokens
4,096 tokens
Context window
128,000 tokens

Transparent token rates

Compare Phi-4-multimodal pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Phi-4-multimodal

Azure Cognitive Services

Official sourceAnnouncement

We are excited to announce Phi-4-multimodal and Phi-4-mini, the newest models in Microsoft’s Phi family of small language models. Learn more.

Azure Cognitive Services

Official sourceAnnouncement

Phi-4-mini brings significant enhancements in multilingual support, reasoning, and mathematics, and now, the long-awaited function calling feature is finally...

Azure Cognitive Services

CoverageRelease Notes

The latest AI models from Phi, 4-mini-instruct and 4-multimodal-instruct, are now available in GitHub Models.

Azure Cognitive Services

Coverage

Microsoft has unveiled two new additions to its Phi-4 family of small language models: Phi-4-multimodal, which integrates speech, vision, and text, and Phi-4-mini., Microsoft has unveiled two new additions to its Phi-4 family of small language models: Phi-4-multimodal, which integrates speech, vision, and text, and Phi

Azure Cognitive Services

CoverageBenchmark

An aggregation page on EmergentMind profiles Phi-4-multimodal-instruct as Microsoft's multimodal LLM in the Phi-4 family, implemented as a compact decoder-only architecture with modality-specific LoRA adapters and specialized encoders, projectors, and routers for efficient instruction tuning across vision-language and speech-language tasks. The page positions Phi-4-multimodal-instruct as a downstream instruction-following instantiation of the Phi-4-Multimodal design, citing the Microsoft Phi-4-Mini Technical Report and noting follow-on research (Ahn et al., September 2025) that uses the model for fine-tuning experiments including pronunciation evaluation. The summary frames Phi-4-multimodal-instruct less as an isolated checkpoint and more as an example of a broader compact multimodal design pattern.

Azure Cognitive Services

Coverage

Microsoft's Phi-4-Mini Technical Report (arXiv:2503.01743v2, March 2025) introduces Phi-4-Multimodal as a compact multimodal model that integrates text, vision, and speech/audio input modalities into a single 3.8-billion-parameter checkpoint. The model uses a novel modality-extension approach built on Mixture-of-LoRAs with modality-specific routers, enabling multiple inference modes — including (vision + language), (vision + speech), and (speech/audio) — without interference between modalities. The speech/audio modality is notably lightweight, with its LoRA component containing only about 460 million parameters, yet Phi-4-Multimodal is reported as ranking first on the OpenASR leaderboard at publication. The report describes Phi-4-Mini's expanded 200K-token vocabulary for multilingual support and group query attention for efficient long-sequence generation, and outlines a separate reasoning-enhanced Phi-4-Mini preview that is not released alongside Phi-4-Mini and Phi-4-Multimodal.

Videos about Phi-4-multimodal

More models around Phi-4-multimodal