Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Phi-4-Mini

Phi-4-Mini is a compact 3.8-billion-parameter dense decoder-only Transformer designed to punch above its weight class. Its architecture incorporates grouped-query attention and shared embeddings to balance efficiency with capability, while a substantially expanded 200K-token vocabulary supports multilingual applications across dozens of languages. The model is engineered for resource-constrained and latency-sensitive environments, making it practical for developers who need capable AI without heavy infrastructure demands.

The model was shaped through a deliberate training recipe combining supervised fine-tuning with direct preference optimization, drawing on curated synthetic data and quality-filtered public web content with a strong emphasis on reasoning-dense material. This lineage helps explain why Phi-4-Mini demonstrates strong performance on math, coding, and multi-step logical tasks—benchmarks show it approaching the quality of models twice its parameter count. It serves as a foundational building block for generative AI features and research into language and multimodal systems, with its design priorities centered on delivering reliable analytical capabilities in settings where memory and compute are at a premium.

Nvidiamicrosoft/phi-4-mini-instructphideprecated

Quick Info

Powered by
Provider
Nvidia
Model key
microsoft/phi-4-mini-instruct
Release date
Dec 1, 2024
Last updated
Sep 5, 2025
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Latest news about Phi-4-Mini

Nvidia

Official sourceRelease Notes

NVIDIA NIM for Large Language Models release 1.12.0 adds "Phi 4 Mini Instruct" as a new language model alongside Gemma 3 1B Instruct, GPT-OSS-20B, GPT-OSS-120B, NVIDIA-Nemotron-Nano-9B-v2, and Stockmark-2-100B-Instruct. The release also introduces support for GB200 and GH200 GPUs and updates CUDA to version 12.9, while The 1.12.0 release notes document several fixed known issues across the 1.12.3 and 1.12.4 patch lines, including GPU requirements and stop-sequence handling for GPT-OSS models, tool-calling support, multi-turn parsing fixes for GPT-OSS-120B, intermittent empty-body 500 errors, and Harmony parser improvements that corre

Videos about Phi-4-Mini

More models around Phi-4-Mini