Phi-4-reasoning-plus is a 14-billion parameter dense decoder-only transformer built to handle sophisticated multi-step reasoning across math, science, and code. The model carries the same foundational architecture as its predecessor Phi-4—a 40-layer transformer with a 5,120 hidden dimension and SwiGLU activation—while producing significantly longer, structured reasoning traces that lay out each step before arriving at a final answer. This design deliberately trades speed for depth, targeting applications where thorough logical decomposition and output quality matter more than quick turnarounds. The model leans into explicit chain-of-thought formatting to make its problem-solving transparent and verifiable, making it especially suited for tasks like competitive math, scientific problem solving, and complex code generation where precision and clarity take priority over latency.
The model undergoes supervised fine-tuning on a curated blend of synthetic chain-of-thought prompts and high-quality public-domain data focused on math, science, and coding skills, with additional alignment data for safety and responsible AI behavior. Reinforcement learning is then applied on top of this foundation, which pushes performance on benchmarks such as AIME, OmniMath, and HumanEvalPlus beyond what Phi-4-reasoning alone achieves. This combination of SFT and RL-trained reasoning produces responses that are roughly 50% longer than the base reasoning variant, reflecting the model's more elaborate internal deliberation. Released as an open-weight model under an MIT license, it provides teams who need transparent, high-precision reasoning without relying on closed API services a capable alternative that can be deployed and fine-tuned locally.