Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Google logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is part of the Gemma 4 family of open models built by Google DeepMind and released under the Apache 2.0 license, with both pre-trained and instruction-tuned variants distributed through Hugging Face. The family spans five sizes (E2B, E4B, 12B, 26B A4B, and 31B) and combines Dense and Mixture-of-Experts architectures, giving developers a range of options for deployment across phones, laptops, and servers. All Gemma 4 models are designed as capable reasoners with configurable thinking modes and maintain multilingual support across over 140 languages.

The 26B A4B variant stands out as the family's MoE configuration, with roughly 25.2 billion total parameters but only about 3.8 to 4 billion activated per token, achieving near the quality of the dense 31B model at substantially lower compute cost. It is built on 30 layers with a 1024-token sliding-window hybrid attention mechanism and supports a 256K-token context window, making it well suited for long-document analysis and agentic workflows. The model accepts text and image input and produces text output, and it is optimized for efficient execution of coding, reasoning, and tool-driven tasks on resource-constrained hardware.

Googlegemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
Google
Model key
gemma-4-26b-a4b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Gemma 4 26B A4B IT

Cortecs

CoverageBenchmark

This Simplismart deployment-focused article targets the Gemma 4 26B MoE variant by name, framing it as delivering near-flagship 31B-level performance at roughly 4B-parameter active inference cost. The piece discusses the MoE routing mechanism, contrasts the 26B-A4B against the 31B dense flagship across reasoning and vi As a third-party provider blog with deployment framing rather than a first-party Google announcement, the article is useful as independent corroboration of the 26B-A4B variant's active-parameter architecture and single-GPU deployability rather than as breaking news. It does not mention Cortecs and is oriented toward Si

Tempr

Coverage

According to a benchr.org launch record published July 28, 2026, Google released the instruction-tuned model with the exact identifier "gemma-4-26b-a4b-it" on April 2, 2026 as part of the Gemma 4 launch. The page records that this exact identifier is made available in AI Studio and via the Gemini API, and frames it as The same benchr record is explicit about what it does and does not publish: the checked release note does not include context-window limits, per-token API pricing, or benchmark results for the exact IT variant. It warns against filling those gaps from sibling Gemma models, older cards, or third-party listings, and trea

Hugging Face

CoverageBenchmark

An independent deep-dive analysis published on kie.ai on July 16, 2026 examines Google's Gemma 4 family as a five-size open-weight lineup under Apache 2.0: E2B (2.3B effective), E4B (4.5B effective), 12B Unified (encoder-free multimodal), 26B-A4B MoE (25.2B total, 3.8B active per token), and 31B dense. The 12B, 26B-A4B The analysis reports that Official Arena AI text rankings place the 31B at Elo 1452 and the 26B-A4B at 1441, both above Gemma 3 27B's 1365, with the 31B scoring 80.0% on LiveCodeBench v6. Community NVFP4 quantizations from Unsloth reportedly deliver 1.5× speedup on NVIDIA Blackwell hardware, fitting the 12B into 11GB V

Hugging Face

CoverageBenchmark

Third-party benchmark aggregator llm-stats.com ranks Gemma 4 26B-A4B at position 117 overall with a composite LLM Stats Score of 22.3 at a blended price of $0.14 per million tokens. The model's capability-tier placement is mixed: it lands in the top half for Data Viz (15 of 64), Audio (19 of 66), Websites (41 of 83), a Per-turn quality analysis shows the model declining from 14.2 at turn 1 to 13.6 at turns 31+, with the Quality Tracker showing +0.82σ improvement in the 7-day window (213 votes) driven by gains in Chat (+1.97σ), SVG (+2.02σ), and Games (+1.96σ), while declining in Websites (-0.64σ) and Audio (-0.63σ). Cost-efficiency c

Google

CoverageBenchmark

Kilo Code's vendor page describes Gemma 4 26B A4B IT as a Google DeepMind instruction-tuned MoE model with a 262,144-token context window, 16,384 max output tokens, and multimodal input modalities covering image, text, and video. It confirms coding-relevant capabilities including function calling, tool choice control, Pricing on the page lists $0.04 per 1M input tokens and $0.22 per 1M output tokens, sourced via OpenRouter, with an example cost of roughly $0.0039 to analyze a 10,000-line codebase at approximately 40k input and 10k output tokens. The Kilo Code integration supports over 500 models across VS Code, JetBrains, CLI, and c

Google

CoverageBenchmark

OpenRouter lists Gemma 4 26B A4B IT as a free-tier hosted variant of Google DeepMind's instruction-tuned MoE model, with 25.2B total parameters and 3.8B active per token, supporting text, image, and video (up to 60s at 1fps) input with a 256K token context window. The page confirms native function calling, configurable Routing is available via OpenRouter's Balanced, Nitro, or Exacto modes across multiple providers including Google AI Studio, SiliconFlow, Parasail, NextBit, and Cloudflare, with a free price tier listed. Structured-output error rates from Google AI Studio average 4.99%, cache hit rates 26.28%, and tool-call error rates

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT