Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Gemma 4 31B

Gemma 4 31B is a 30.7-billion-parameter dense multimodal model that accepts text and images and produces text. It is designed for demanding language-and-vision work, including document understanding, coding, reasoning, and native function calling, with configurable thinking or reasoning behavior and multilingual support across more than 140 languages. Its practical appeal is the combination of strong general-purpose capabilities with a large context window, making it suitable for complex document workflows, code-related tasks, and applications that need tool-driven responses.

Built from Gemini 3 research and technology, the model emphasizes a high level of intelligence relative to its parameter count. The 31B configuration is intended to offer more capability than smaller Gemma 4 variants, while a separate 2B-and-4B line targets maximum efficiency on mobile and IoT devices. Community deployment testing has also exercised the 31B model with FP8 quantization, multi-user concurrency, tool calling, and multimodal vision, indicating that it can be deployed in specialized local or dedicated infrastructure.

NovitaAIgoogle/gemma-4-31b-itgemma

Quick Info

Powered by
Provider
NovitaAI
Model key
google/gemma-4-31b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Gemma 4 31B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 31B

NovitaAI

CoverageBenchmark

According to a recent LinkedIn post from FriendliAI, the company is emphasizing the availability of Gemma-4-31B-it on its Model APIs and Dedicated Endpoints, positi...

NovitaAI

CoverageBenchmark

Alibaba's new open-source Qwen3.6-35B-A3B activates just three of its 35 billion parameters at a time, yet beats Google's larger Gemma 4-31B on coding and reasoning benchmarks.

NovitaAI

CoverageBenchmark

Gemma 4 31B dense benchmarks on DGX Spark with runtime FP8 quantization. Includes single-node (TP=1), dual-node (TP=2), multi-user concurrency, tool calling validation, and multimodal vision. Getting Gemma 4 running on v…

NovitaAI

CoverageBenchmark

OpenRouter's provider routing page for google/gemma-4-31b-it documents NovitaAI as one of the paid hosts for the model, listing $0.14 per 1M input tokens and $0.40 per 1M output tokens with no cache price tier, a P50 latency of 1.29 seconds, throughput of 6 tokens per second, and 89.88% uptime — figures that place Novi The same OpenRouter entry confirms the underlying Gemma 4 31B Instruct specification available through NovitaAI: a 30.7B dense multimodal model from Google DeepMind with text-and-image input and text output, a 262K token context window, configurable thinking/reasoning mode, native function calling, multilingual support

Videos about Gemma 4 31B

More models around Gemma 4 31B