NovitaAI
According to a recent LinkedIn post from FriendliAI, the company is emphasizing the availability of Gemma-4-31B-it on its Model APIs and Dedicated Endpoints, positi...
Model details
Gemma 4 31B is a 30.7-billion-parameter dense multimodal model that accepts text and images and produces text. It is designed for demanding language-and-vision work, including document understanding, coding, reasoning, and native function calling, with configurable thinking or reasoning behavior and multilingual support across more than 140 languages. Its practical appeal is the combination of strong general-purpose capabilities with a large context window, making it suitable for complex document workflows, code-related tasks, and applications that need tool-driven responses.
Built from Gemini 3 research and technology, the model emphasizes a high level of intelligence relative to its parameter count. The 31B configuration is intended to offer more capability than smaller Gemma 4 variants, while a separate 2B-and-4B line targets maximum efficiency on mobile and IoT devices. Community deployment testing has also exercised the 31B model with FP8 quantization, multi-user concurrency, tool calling, and multimodal vision, indicating that it can be deployed in specialized local or dedicated infrastructure.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
NovitaAI
According to a recent LinkedIn post from FriendliAI, the company is emphasizing the availability of Gemma-4-31B-it on its Model APIs and Dedicated Endpoints, positi...
NovitaAI
Alibaba's new open-source Qwen3.6-35B-A3B activates just three of its 35 billion parameters at a time, yet beats Google's larger Gemma 4-31B on coding and reasoning benchmarks.
NovitaAI
Gemma 4 31B dense benchmarks on DGX Spark with runtime FP8 quantization. Includes single-node (TP=1), dual-node (TP=2), multi-user concurrency, tool calling validation, and multimodal vision. Getting Gemma 4 running on v…
NovitaAI
OpenRouter's provider routing page for google/gemma-4-31b-it documents NovitaAI as one of the paid hosts for the model, listing $0.14 per 1M input tokens and $0.40 per 1M output tokens with no cache price tier, a P50 latency of 1.29 seconds, throughput of 6 tokens per second, and 89.88% uptime — figures that place Novi The same OpenRouter entry confirms the underlying Gemma 4 31B Instruct specification available through NovitaAI: a 30.7B dense multimodal model from Google DeepMind with text-and-image input and text output, a 262K token context window, configurable thinking/reasoning mode, native function calling, multilingual support
This exact model name is also listed by 3 other providers.