Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Gemma 4 26B A4B

Gemma 4 26B A4B is a Mixture-of-Experts model from Google's DeepMind team that brings advanced reasoning capabilities to a wide range of deployment scenarios. With 26 billion total parameters and roughly 4 billion active parameters per forward pass, the architecture enables powerful capabilities while managing computational demands. The model is designed as a highly capable reasoner with configurable thinking modes, supports native vision for processing images with variable aspect ratios and resolutions, and handles function-calling and structured JSON output—making it well-suited for complex tasks like text generation, coding, and multi-step reasoning in professional and developer workflows.

The Gemma 4 family builds on Gemini 3 research, and this particular variant carries forward that lineage into an open-weight format accessible to developers and researchers. Multilingual support spans over 140 languages, and the extensive context window of around 262K tokens enables long-form document understanding and complex multi-turn conversations. By offering open-weights models in instruction-tuned variants alongside dense and MoE architectures across multiple sizes, Google DeepMind has designed the family for scalability—from high-end mobile devices to server deployments. This accessibility, combined with the model's strength in structured outputs and tool use, positions it for both exploratory research and production applications where control over model behavior matters.

NovitaAIgoogle/gemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
NovitaAI
Model key
google/gemma-4-26b-a4b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.13
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Gemma 4 26B A4B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 26B A4B

NovitaAI

CoverageBenchmark

Simplismart published a technical analysis on August 30, 2026, focused specifically on Google's Gemma 4 26B A4B Mixture-of-Experts model. The piece details the model's "Active 4B" MoE routing design, in which only around 4 billion of the model's 26 billion total parameters activate per forward pass, delivering near-fla The article benchmarks the 26B A4B against the 31B dense variant, with particular strength in image, OCR, and document understanding tasks. It also documents that the model can be deployed and served on a single NVIDIA A100, providing official sampling parameters and a recommended deployment configuration. The comparis

Videos about Gemma 4 26B A4B

More models around Gemma 4 26B A4B