Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Gemma-4-31B-IT

As the dense flagship of Google DeepMind's open Gemma 4 family, the 31B instruction-tuned variant is built around a 30.7-billion-parameter architecture paired with Multi-Token Prediction drafter heads that enable speculative decoding for faster inference. It handles text, image, and video inputs while producing text outputs, and pairs that multimodality with the cataloged API limit context window and multilingual coverage spanning more than 140 languages. The Apache 2.0 license keeps the weights and surrounding tooling broadly accessible, sitting alongside sibling configurations in the family that range from efficient E2B and E4B on-device variants to a 26B-A4B mixture-of-experts model, allowing the 31B to serve as the dense reference point for downstream deployments.

In practice the model is aimed at teams that need a single open checkpoint for coding assistants, document understanding, and reasoning workflows, with configurable thinking modes and native function calling supporting tool-use and structured JSON at the model level rather than through wrappers. The release landed on April 2, 2026, and the community has already pushed NVFP4 quantizations that run roughly 1.5x faster on Nvidia Blackwell silicon, with hosted providers such as Cerebras reporting throughput above 1,500 tokens per second for the 31B. Multiple routing providers—including Google AI Studio, SambaNova, Together, DeepInfra, and Crusoe—offer access with reported GPQA Diamond scores ranging from 69.2% to 85.7%, giving integrators room to match cost, latency, and tool-calling reliability to their workload.

Nvidiagoogle/gemma-4-31b-itgemma

Quick Info

Powered by
Provider
Nvidia
Model key
google/gemma-4-31b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
16,384 tokens
Context window
256,000 tokens

Latest news about Gemma-4-31B-IT

Nvidia

CoverageBenchmark

A Googler-authored Google Cloud Community article dated August 14, 2026 provides a detailed technical guide on serving MTP (Multi-Token Prediction)-based Gemma 4 models for inference performance. It covers empirical benchmarks comparing no-MTP vs. k=1 vs. k=2 speculative decoding, the rationale for using MIG and vGPU s The article also details how Gemma 4's MTP speculative decoding saves VRAM and KV-cache compared to previous approaches through a shared KV-cache pool and single-pass parallel verification flow, provides fine-tuning guidance to preserve MTP proposer alignment, lists a production observability checklist, and includes a

Nvidia

CoverageBenchmark

An independent July 16, 2026 deep-dive blog analyzes the Gemma 4 family's five-size open-weight lineup under Apache 2.0: E2B, E4B, 12B Unified, 26B-A4B MoE, and 31B dense. The 31B features a 256K context window, native multimodality across text/image/audio/video, Multi-Token Prediction drafter heads for speculative dec The piece also reports that community NVFP4 quants from Unsloth run approximately 1.5× faster on Nvidia Blackwell hardware, with the 12B NVFP4 fitting into 11GB VRAM. Cerebras is cited as serving the 31B at over 1,500 tokens per second. Native tool calling and structured JSON operate at the model level rather than thro

Nvidia

CoverageBenchmark

The OpenRouter routing page for the free variant of Gemma 4 31B Instruct confirms the model's metadata: a 30.7B dense multimodal model supporting text and image input with text output, a 256K (262K displayed) token context window, configurable thinking/reasoning mode, native function calling, and multilingual support a The page lists providers routing the model including Google AI Studio, SambaNova, Together, DeepInfra, and Crusoe, with GPQA Diamond scores ranging from 69.2% to 85.7% across providers. Tool-call error rate on Google AI Studio averages 2.43%, and availability over the three-day window was 99.40%. Notably, the provider

Videos about Gemma-4-31B-IT

More models around Gemma-4-31B-IT