Currently listed through these providers:
Model details
Gemma 4 31B IT
Gemma 4 31B IT is the instruction-tuned flagship of Google DeepMind's open Gemma 4 family, distilled from Gemini 3 research and built around a 30.7-billion parameter dense architecture. It is released under Apache 2.0, paired with a pre-trained sibling, and joins sibling variants that span dense and mixture-of-experts designs from E2B up through 26B A4B and 31B. The family targets text generation, coding, and reasoning workloads across deployment environments ranging from phones to servers, with the 31B Instruct variant tuned for chat, task completion, and agentic workflows.
The model processes text and image inputs into text outputs, drawing on a 256K-token context window with hybrid attention for efficient long-context handling, and accepts video on selected variants. It supports a configurable thinking mode, native function calling, structured output, and multilingual understanding across more than 140 languages. Reported strong suits include reasoning benchmarks such as GPQA Diamond at 85.7%, alongside coding and document understanding tasks, making it well suited for coding assistants, multimodal analysis, document Q&A, and tool-driven agents that benefit from an open-weights deployment path.
Quick Info
Powered by- Provider
- ai&
- Model key
- google/gemma-4-31b-it
- Release date
- Apr 2, 2026
- Last updated
- Apr 2, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.50
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Gemma 4 31B IT pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemma 4 31B IT
No articles yet. Fetch the latest news to show it here.
Videos about Gemma 4 31B IT
More models around Gemma 4 31B IT
This exact model name is also listed by 23 other providers.