Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

Gemma 4 26B A4B

Google DeepMind's Gemma 4 26B A4B is the first Mixture-of-Experts entry in the Gemma family, pairing 25.2 billion total parameters with a selective routing design that activates roughly 3.8 billion parameters per token. A router picks 8 of 128 experts plus a shared expert for every token, so generation behaves computationally like a 4B model while the full capacity remains loaded for routing. The model shares the Gemma 4 backbone with its dense siblings, including a hybrid local-plus-global attention scheme, unified keys and values on global layers, and proportional RoPE positional encoding. Native system prompts and a configurable thinking mode triggered by a leading control token give developers explicit control over step-by-step reasoning before final answers are produced.

In Google's published benchmark comparisons, the MoE variant trails the denser 31B sibling by only about 1 to 5 percentage points on core reasoning, math, and coding evaluations such as MMLU Pro, GPQA Diamond, LiveCodeBench, and AIME 2026, while delivering major gains over the previous Gemma 3 generation. Vision and document understanding benchmarks like MMMU Pro, MATH-Vision, and OmniDocBench 1.5 stay close to the dense flagship, and the 256K-token context window supports deep document analysis and RAG-style workloads. Open weights under Apache 2.0, single-A100 BF16 viability, quantised checkpoints that fit consumer GPUs, and native function calling make the model well suited to chatbots, coding assistants, agentic tool use, and high-concurrency production deployments where token cost and latency matter more than squeezing out the last few accuracy points.

CoreWeavegoogle/gemma-4-26B-A4B-itgemma

Quick Info

Powered by
Provider
CoreWeave
Model key
google/gemma-4-26B-A4B-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Gemma 4 26B A4B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 26B A4B

CoreWeave

CoverageBenchmark

Simplismart's comparison is explicitly titled around the Gemma 4 26B A4B MoE variant, framing it as delivering near-flagship reasoning at roughly 4B-inference cost. The page walks through the model's MoE routing design, training data and safety standards, and positions the 26B A4B against the 31B dense sibling in a dec The article further details the hardware requirements and throughput advantage of running the 26B A4B on a single A100, including official sampling and deployment configuration guidance for developers reproducing results. It also documents the shared architecture backbone across the Gemma 4 family that the MoE variant

Videos about Gemma 4 26B A4B

More models around Gemma 4 26B A4B