As the dense flagship of Google DeepMind's open Gemma 4 family, the 31B instruction-tuned variant is built around a 30.7-billion-parameter architecture paired with Multi-Token Prediction drafter heads that enable speculative decoding for faster inference. It handles text, image, and video inputs while producing text outputs, and pairs that multimodality with the cataloged API limit context window and multilingual coverage spanning more than 140 languages. The Apache 2.0 license keeps the weights and surrounding tooling broadly accessible, sitting alongside sibling configurations in the family that range from efficient E2B and E4B on-device variants to a 26B-A4B mixture-of-experts model, allowing the 31B to serve as the dense reference point for downstream deployments.
In practice the model is aimed at teams that need a single open checkpoint for coding assistants, document understanding, and reasoning workflows, with configurable thinking modes and native function calling supporting tool-use and structured JSON at the model level rather than through wrappers. The release landed on April 2, 2026, and the community has already pushed NVFP4 quantizations that run roughly 1.5x faster on Nvidia Blackwell silicon, with hosted providers such as Cerebras reporting throughput above 1,500 tokens per second for the 31B. Multiple routing providers—including Google AI Studio, SambaNova, Together, DeepInfra, and Crusoe—offer access with reported GPQA Diamond scores ranging from 69.2% to 85.7%, giving integrators room to match cost, latency, and tool-calling reliability to their workload.