Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Charm Hyper logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is part of the Gemma 4 family of open models built by Google DeepMind and released under the Apache 2.0 license, with both pre-trained and instruction-tuned variants distributed through Hugging Face. The family spans five sizes (E2B, E4B, 12B, 26B A4B, and 31B) and combines Dense and Mixture-of-Experts architectures, giving developers a range of options for deployment across phones, laptops, and servers. All Gemma 4 models are designed as capable reasoners with configurable thinking modes and maintain multilingual support across over 140 languages.

The 26B A4B variant stands out as the family's MoE configuration, with roughly 25.2 billion total parameters but only about 3.8 to 4 billion activated per token, achieving near the quality of the dense 31B model at substantially lower compute cost. It is built on 30 layers with a 1024-token sliding-window hybrid attention mechanism and supports a 256K-token context window, making it well suited for long-document analysis and agentic workflows. The model accepts text and image input and produces text output, and it is optimized for efficient execution of coding, reasoning, and tool-driven tasks on resource-constrained hardware.

Charm Hypergemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
Charm Hyper
Model key
gemma-4-26b-a4b-it
Release date
Apr 30, 2026
Last updated
Jul 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.098
Output token cost
$0.334

Limits

Output tokens
25,600 tokens
Context window
256,000 tokens

Transparent token rates

Compare Gemma 4 26B A4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 26B A4B IT

Cortecs

CoverageBenchmark

This Simplismart deployment-focused article targets the Gemma 4 26B MoE variant by name, framing it as delivering near-flagship 31B-level performance at roughly 4B-parameter active inference cost. The piece discusses the MoE routing mechanism, contrasts the 26B-A4B against the 31B dense flagship across reasoning and vi As a third-party provider blog with deployment framing rather than a first-party Google announcement, the article is useful as independent corroboration of the 26B-A4B variant's active-parameter architecture and single-GPU deployability rather than as breaking news. It does not mention Cortecs and is oriented toward Si

Tempr

Coverage

According to a benchr.org launch record published July 28, 2026, Google released the instruction-tuned model with the exact identifier "gemma-4-26b-a4b-it" on April 2, 2026 as part of the Gemma 4 launch. The page records that this exact identifier is made available in AI Studio and via the Gemini API, and frames it as The same benchr record is explicit about what it does and does not publish: the checked release note does not include context-window limits, per-token API pricing, or benchmark results for the exact IT variant. It warns against filling those gaps from sibling Gemma models, older cards, or third-party listings, and trea

Hugging Face

CoverageBenchmark

An independent deep-dive analysis published on kie.ai on July 16, 2026 examines Google's Gemma 4 family as a five-size open-weight lineup under Apache 2.0: E2B (2.3B effective), E4B (4.5B effective), 12B Unified (encoder-free multimodal), 26B-A4B MoE (25.2B total, 3.8B active per token), and 31B dense. The 12B, 26B-A4B The analysis reports that Official Arena AI text rankings place the 31B at Elo 1452 and the 26B-A4B at 1441, both above Gemma 3 27B's 1365, with the 31B scoring 80.0% on LiveCodeBench v6. Community NVFP4 quantizations from Unsloth reportedly deliver 1.5× speedup on NVIDIA Blackwell hardware, fitting the 12B into 11GB V

Hugging Face

CoverageBenchmark

Third-party benchmark aggregator llm-stats.com ranks Gemma 4 26B-A4B at position 117 overall with a composite LLM Stats Score of 22.3 at a blended price of $0.14 per million tokens. The model's capability-tier placement is mixed: it lands in the top half for Data Viz (15 of 64), Audio (19 of 66), Websites (41 of 83), a Per-turn quality analysis shows the model declining from 14.2 at turn 1 to 13.6 at turns 31+, with the Quality Tracker showing +0.82σ improvement in the 7-day window (213 votes) driven by gains in Chat (+1.97σ), SVG (+2.02σ), and Games (+1.96σ), while declining in Websites (-0.64σ) and Audio (-0.63σ). Cost-efficiency c

Charm Hyper

CoverageBenchmark

The OpenRouter product page describes Gemma 4 26B A4B IT as a Google DeepMind instruction-tuned Mixture-of-Experts model with 25.2B total parameters but only 3.8B activated per token, delivering near-31B quality at lower compute cost. The excerpt confirms a 256K-token context window, multimodal input (text, images, and OpenRouter's free-tier listing reports real-world performance: P50 latency of 0.86 seconds and throughput of 34 tokens per second against Google AI Studio, with 99.41% uptime over the tracked window. GPQA Diamond scores range from 72.0% (auto-routing) to 77.2% (SiliconFlow), and TAU-Bench scores reach 69.5% across prov

Charm Hyper

CoverageBenchmark

The ModelBench catalog lists Gemma 4 26B A4B IT as an open-weights Google instruction model released on April 2, 2026, and explicitly names Charm Hyper among 21 inference providers hosting it. According to the supplied excerpt, Charm Hyper serves the model under the provider-specific ID "gemma-4-26b-a4b-it" with a 256K Across the same ModelBench page, alternative hosts show pricing ranging from $0.042/M input at Kilo Gateway up to $0.13/M input at providers such as Hugging Face and NovitaAI, placing Charm Hyper's $0.116 input rate in the upper-mid tier while its $0.38 output rate sits below the typical $0.40 ceiling. Most providers c

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT