Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Gemma 4 31B IT

Gemma 4 31B IT is the largest model in Google DeepMind's Gemma 4 open-weight family, shaped by Gemini 3 research and engineered to push intelligence-per-parameter further than earlier Gemma releases. It is a dense, instruction-tuned, vision-language model, meaning it accepts both text and image inputs, has been post-trained to follow instructions, and processes visual information natively rather than through a bolted-on adapter. This design lineage, drawing directly from frontier Gemini 3 research, signals that the model is intended as a capable general assistant rather than a narrowly specialized tool, while still remaining within the open-weights ecosystem that has defined the Gemma family.

In practical terms, Gemma 4 31B IT is positioned for workloads that demand both reasoning and visual understanding, including coding assistance, agentic workflows that chain tool calls together, structured document extraction, and visual question answering over images or screenshots. Reported benchmark performance is strong for an open model of its size, with approximately 2,150 Codeforces ELO reflecting competitive coding ability, 89.2% on AIME 2026 indicating advanced mathematical reasoning, 84.3% on GPQA Diamond suggesting graduate-level science understanding, and 76.9% on MMMU Pro showing robust multimodal comprehension. These strengths make the model a fitting choice for developers building assistive agents, document understanding pipelines, or reasoning-heavy applications where open deployment and strong baseline performance matter more than absolute frontier scores.

Requestygemma-4-31b-itgemma

Quick Info

Powered by
Provider
Requesty
Model key
gemma-4-31b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
262,144 tokens

Latest news about Gemma 4 31B IT

Requesty

Official sourceBenchmark

Requesty exposes DeepInfra's flex tier of Gemma 4 31B IT as deepinfra/google/gemma-4-31B-it:flex, priced at $0.10/1M input and $0.30/1M output (2.9x input ratio) with a 262K-token context window and no published max output, US-served with no data retention, not used for training, added April 2026, and ZDR. The page adv For developers, the DeepInfra :flex variant is the cheapest paid Requesty-routed path to Gemma 4 31B IT weights as of September 4, 2026, undercutting DeepInfra's own standard tier ($0.13/$0.38) and far below Sail Research's $0.40/$0.60, while preserving the model's multimodal text+image behavior and the 262K context. T

ai&

Coverage

A July 13, 2026 Google Cloud Community article by Vipul Raja provides a deep technical look at Gemma 4 31B IT, explicitly naming the instruction-tuned 31B variant released by Google DeepMind in April 2026 as part of the Gemma 4 family. The piece breaks down the naming convention (Gemma family, 4th generation, 31 billio The same article details key capabilities directly tied to the 31B IT variant: a configurable "thinking" mode for step-by-step problem-solving on complex tasks, native support for function calling and structured output suited to agentic workflows, and an Apache 2.0 license that allows developers and researchers to free

Requesty

Official sourceBenchmark

Requesty lists Google's first-party Gemma 4 31B IT endpoint at the model id google/gemma-4-31b-it, priced free for both input and output with no cache read/write pricing, a 262K-token context window, an 8K-token max output, and an April 2026 add date. The deployment is served globally, data retention is flagged Yes whi The useful developer signal is that Gemma 4 31B IT remains callable at zero cost through Google's own Gemini API endpoint via Requesty's OpenAI-compatible router, using google/gemma-4-31b-it as the model field with no per-token charges for input, output, or (where applicable) cache reads. The trade-off versus paid endp

Requesty

CoverageBenchmark

A technical deep dive on Gemma 4 describes the family as multimodal, handling text and image input (with audio on small models) and generating text output, with a context window of up to 256K tokens and multilingual support across 140+ languages. The release includes both pre-trained and instruction-tuned variants. The Google DeepMind's stated positioning is "unprecedented intelligence-per-parameter" purpose-built for advanced reasoning and agentic workflows. Since the first Gemma generation, the ecosystem has seen over 400 million downloads and more than 100,000 community variants. The article frames Gemma 4 against DeepSeek R2, Qwe

Merge Gateway

Coverage

Google Cloud announced Gemma 4 availability on Vertex AI on April 2, 2026, positioning it as the company's "most capable family of open models" built from the same research as Gemini 3 and released under an Apache 2.0 license. The blog explicitly names the Gemma 4 31B dense model as a variant suited for "complex enterp The post also notes that the Gemma 4 26B MoE variant will become fully managed and serverless on Model Garden "over the coming days," and highlights integration with the Agent Development Kit (ADK) for building and deploying AI agents. Enterprise-focused features include deployment across Sovereign Cloud solutions for

Requesty

Official sourceBenchmark

Requesty lists DeepInfra's standard deployment of Gemma 4 31B IT at deepinfra/google/gemma-4-31B-it, priced at $0.13/1M input and $0.38/1M output (2.9x input ratio), with a 262K-token context window, US-served, no data retention, not used for training, ZDR, added April 2026, and 4/8 capabilities including vision, reaso The developer takeaway is that DeepInfra's standard Gemma 4 31B IT endpoint is a mid-priced option on the Requesty router, sitting between the cheaper DeepInfra :flex flex tier ($0.10/$0.30) and the premium Sail Research tier ($0.40/$0.60), while Google's own Gemini API endpoint remains free. Because all three paid tie

Requesty

Official sourceBenchmark

Requesty's model-overview page consolidates Gemma 4 31B IT, described as Google DeepMind's 30.7B dense multimodal model with text+image input, text output, a 256K-token context window, configurable reasoning mode, native function calling, multilingual support across 140+ languages, and an Apache 2.0 license, noting str The developer-relevant change is the availability of a managed, auto-routing Requesty id for Gemma 4 31B IT that masks provider differences, so a single gemma-4-31b-it string picks the cheapest, healthy endpoint at call time and survives provider changes underneath. With Google Gemini API offered free and DeepInfra sta

Requesty

Official sourceBenchmark

Requesty now routes Gemma 4 31B IT through a new Sail Research Co. endpoint at the model id sail/gemma-4-31b-it, priced at $0.40/1M input, $0.60/1M output, and $0.20/1M for cached reads, with a 262K-token context window and a 66K-token max output. The deployment is US-served, marked as no data retention and not used fo For developers, the practical signal is that Gemma 4 31B IT (Google DeepMind's 30.7B dense multimodal model with image+text input and text output, Apache 2.0 licensed) is available on the Requesty router from three independent providers with a 1.3x price spread: Google Gemini API (free, 8K output, data retained but no-

Videos about Gemma 4 31B IT

More models around Gemma 4 31B IT