Google DeepMind's Gemma 4 26B A4B is the first Mixture-of-Experts entry in the Gemma family, pairing 25.2 billion total parameters with a selective routing design that activates roughly 3.8 billion parameters per token. A router picks 8 of 128 experts plus a shared expert for every token, so generation behaves computationally like a 4B model while the full capacity remains loaded for routing. The model shares the Gemma 4 backbone with its dense siblings, including a hybrid local-plus-global attention scheme, unified keys and values on global layers, and proportional RoPE positional encoding. Native system prompts and a configurable thinking mode triggered by a leading control token give developers explicit control over step-by-step reasoning before final answers are produced.
In Google's published benchmark comparisons, the MoE variant trails the denser 31B sibling by only about 1 to 5 percentage points on core reasoning, math, and coding evaluations such as MMLU Pro, GPQA Diamond, LiveCodeBench, and AIME 2026, while delivering major gains over the previous Gemma 3 generation. Vision and document understanding benchmarks like MMMU Pro, MATH-Vision, and OmniDocBench 1.5 stay close to the dense flagship, and the 256K-token context window supports deep document analysis and RAG-style workloads. Open weights under Apache 2.0, single-A100 BF16 viability, quantised checkpoints that fit consumer GPUs, and native function calling make the model well suited to chatbots, coding assistants, agentic tool use, and high-concurrency production deployments where token cost and latency matter more than squeezing out the last few accuracy points.