Currently listed through these providers:
Model details
Gemma 4 31B
Gemma 4 31B is part of the open model family built by Google DeepMind and described as their most intelligent open release, built from Gemini 3 research to maximize intelligence-per-parameter. The family spans five sizes ranging from efficient E2B and E4B variants through 12B, 26B A4B, and 31B, giving deployers options that scale from phones and laptops up to server-class hardware. Across the lineup, Gemma 4 combines dense and mixture-of-experts designs, which helps the larger variants like 31B balance raw capability against efficiency for production serving.
The 31B variant is multimodal at input, accepting text and images while producing text outputs, and ships with open weights under the Apache 2.0 license in both pre-trained and instruction-tuned forms. A context window of up to 256K tokens and broad multilingual coverage across more than 140 languages make it well suited to long-document reasoning, code generation, and general text tasks where extended context and language variety matter. Configurable thinking modes, native image handling, and the Gemini 3 research lineage position it as a strong general-purpose choice for teams standardizing on an open-weight model that still benefits from frontier-model techniques.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- TEE/gemma4-31b
- Release date
- Apr 4, 2026
- Last updated
- Apr 4, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $1.00
Limits
- Input tokens
- 262,144 tokens
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Gemma 4 31B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.