Gemini 3.1 Pro Preview builds on the multimodal Gemini 3 series as an improved checkpoint rather than a new architecture, retaining the 1M-token that quick-info value window while extending reasoning quality across text, image, video, audio, and code inputs. According to OpenRouter, the 3.1 update introduces measurable gains in software engineering benchmarks and real-world coding environments, alongside stronger autonomous task execution in structured domains such as finance and spreadsheet workflows. The model adds a medium thinking level to help balance cost, speed, and performance for complex multi-step agentic pipelines, and Google documented targeted improvements in instruction following on agentic tasks, front-end code generation, structured output fidelity, and video understanding weights trained on a wider content distribution.
In benchmark comparisons against Gemini 3 Pro, WhatLLM.org reports gains of +4.2 percentage points on SWE-Bench Verified and +5.2 percentage points on AIME 2025, with an LM Arena ELO of 1489. The same source notes Gemini 3.1 Pro Preview is priced roughly twelve times lower than Claude Opus 4.5. Google Research subsequently used Gemini 3.1 Pro to power Gemini-SQL2, a text-to-SQL system that reached 80.04% execution accuracy on the BIRD Text-to-SQL Leaderboard's Single Model track, surpassing Google's prior Gemini-SQL entry. These results position the model as well-suited to long-horizon development tasks, agentic orchestration, and high-context enterprise automation where stable tool use and multimodal reasoning matter more than speculative capability claims.