Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is part of the Gemma 4 family of open models built by Google DeepMind and released under the Apache 2.0 license, with both pre-trained and instruction-tuned variants distributed through Hugging Face. The family spans five sizes (E2B, E4B, 12B, 26B A4B, and 31B) and combines Dense and Mixture-of-Experts architectures, giving developers a range of options for deployment across phones, laptops, and servers. All Gemma 4 models are designed as capable reasoners with configurable thinking modes and maintain multilingual support across over 140 languages.

The 26B A4B variant stands out as the family's MoE configuration, with roughly 25.2 billion total parameters but only about 3.8 to 4 billion activated per token, achieving near the quality of the dense 31B model at substantially lower compute cost. It is built on 30 layers with a 1024-token sliding-window hybrid attention mechanism and supports a 256K-token context window, making it well suited for long-document analysis and agentic workflows. The model accepts text and image input and produces text output, and it is optimized for efficient execution of coding, reasoning, and tool-driven tasks on resource-constrained hardware.

Requestygemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
Requesty
Model key
gemma-4-26b-a4b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.34

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Gemma 4 26B A4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 26B A4B IT

Cortecs

CoverageBenchmark

This Simplismart deployment-focused article targets the Gemma 4 26B MoE variant by name, framing it as delivering near-flagship 31B-level performance at roughly 4B-parameter active inference cost. The piece discusses the MoE routing mechanism, contrasts the 26B-A4B against the 31B dense flagship across reasoning and vi As a third-party provider blog with deployment framing rather than a first-party Google announcement, the article is useful as independent corroboration of the 26B-A4B variant's active-parameter architecture and single-GPU deployability rather than as breaking news. It does not mention Cortecs and is oriented toward Si

Requesty

Official sourceBenchmark

The DeepInfra Flex deployment of Gemma 4 26B A4B IT is priced at $0.06 per 1M input and $0.27 per 1M output tokens (4.9x input:output ratio), making it the cheapest of the four endpoints. It offers a 262K token context window, chat API, US serving region, no data retention for training, and 4/8 capability flags includi Model id `deepinfra/google/gemma-4-26B-A4B-it:flex` was added in April 2026 and is reached through the Requesty router at https://router.requesty.ai/v1 with no additional routing or failover. Workload estimates at these rates run from $0.0083 for 100K input + 10K output up to $0.83 for 10M input + 1M output.

Tempr

Coverage

According to a benchr.org launch record published July 28, 2026, Google released the instruction-tuned model with the exact identifier "gemma-4-26b-a4b-it" on April 2, 2026 as part of the Gemma 4 launch. The page records that this exact identifier is made available in AI Studio and via the Gemini API, and frames it as The same benchr record is explicit about what it does and does not publish: the checked release note does not include context-window limits, per-token API pricing, or benchmark results for the exact IT variant. It warns against filling those gaps from sibling Gemma models, older cards, or third-party listings, and trea

Hugging Face

CoverageBenchmark

An independent deep-dive analysis published on kie.ai on July 16, 2026 examines Google's Gemma 4 family as a five-size open-weight lineup under Apache 2.0: E2B (2.3B effective), E4B (4.5B effective), 12B Unified (encoder-free multimodal), 26B-A4B MoE (25.2B total, 3.8B active per token), and 31B dense. The 12B, 26B-A4B The analysis reports that Official Arena AI text rankings place the 31B at Elo 1452 and the 26B-A4B at 1441, both above Gemma 3 27B's 1365, with the 31B scoring 80.0% on LiveCodeBench v6. Community NVFP4 quantizations from Unsloth reportedly deliver 1.5× speedup on NVIDIA Blackwell hardware, fitting the 12B into 11GB V

Hugging Face

CoverageBenchmark

Third-party benchmark aggregator llm-stats.com ranks Gemma 4 26B-A4B at position 117 overall with a composite LLM Stats Score of 22.3 at a blended price of $0.14 per million tokens. The model's capability-tier placement is mixed: it lands in the top half for Data Viz (15 of 64), Audio (19 of 66), Websites (41 of 83), a Per-turn quality analysis shows the model declining from 14.2 at turn 1 to 13.6 at turns 31+, with the Quality Tracker showing +0.82σ improvement in the 7-day window (213 votes) driven by gains in Chat (+1.97σ), SVG (+2.02σ), and Games (+1.96σ), while declining in Websites (-0.64σ) and Audio (-0.63σ). Cost-efficiency c

Requesty

Official sourceBenchmark

Novita AI's deployment of Gemma 4 26B A4B IT is priced at $0.13 per 1M input and $0.40 per 1M output (3.1x ratio), with $0.13 per 1M cache read, and offers a 262K token context with 131K max output — the largest output window of the available endpoints. It was added in April 2026, served from the US, with no data retai The endpoint reports 3/8 capability flags (Vision, Reasoning, Tool calling) — fewer than the aggregator's 5/8 — suggesting some capabilities listed at the model level may not be available on this specific provider. Accessed via model id `novita/google/gemma-4-26b-a4b-it` through Requesty's OpenAI-compatible router, wit

Requesty

Official sourceBenchmark

The standard DeepInfra deployment of Gemma 4 26B A4B IT is priced at $0.07 per 1M input and $0.34 per 1M output, with $0.07 per 1M cache read pricing — a feature the Flex tier lacks. It shares the 262K context window, chat API, US region, April 2026 addition date, and no-training data retention with sibling endpoints. Accessible via model id `deepinfra/google/gemma-4-26B-A4B-it` through the Requesty router with OpenAI-compatible calls. The page notes no benchmarks are published for this exact variant, and workload cost estimates at these rates range from $0.0104 (100K+10K) to $1.04 (10M+1M tokens).

Requesty

Official sourceBenchmark

Requesty's aggregator page for Gemma 4 26B A4B IT describes it as a Mixture of Experts model from Google DeepMind with a 262K token context window, supporting text and image input, native thinking mode, function calling, structured output, and 140+ languages. The page lists it as an open-weights model with capabilities The model is served across 4 endpoints from 3 providers — DeepInfra (two tiers at $0.06 and $0.07 per 1M input tokens), Novita AI ($0.13), and Parasail ($0.13) — with output pricing ranging from $0.27 to $0.40 per 1M tokens, a 2.3x price spread, and 131K max output on some endpoints. Requesty's managed id `gemma-4-26b-

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT