Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Amazon Bedrock logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is an instruction-tuned model from Google DeepMind designed as part of the open Gemma 4 family, with weights published on Hugging Face under an Apache 2.0 license. It uses a Mixture-of-Experts architecture with roughly 25.2B total parameters but only about 3.8B activated per token, an efficiency-oriented design that aims to deliver denser-model quality at a lower per-token compute cost. This sparse-activation approach makes it well suited for developers who want a capable open-weight model that can run with lighter inference budgets than comparable dense checkpoints.

The model is built for multimodal assistants and agent-style applications: it accepts text, image, and video input (including short video clips) while producing text output, and it supports native function calling, configurable thinking or reasoning mode, and structured outputs. A long context window makes it practical for document-heavy or multi-turn reasoning workflows, and the open-weight release lets teams fine-tune, self-host, or integrate the model into existing pipelines. Practically, it fits use cases such as grounded visual question answering, tool-using chatbots, and reasoning-heavy text tasks where an open, Apache-licensed Google model is preferred over proprietary alternatives.

Amazon Bedrockgoogle.gemma-4-26b-a4bgemma

Quick Info

Powered by
Provider
Amazon Bedrock
Model key
google.gemma-4-26b-a4b
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
AI SDK package
@ai-sdk/amazon-bedrock/mantle
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.13
Output token cost
$0.40

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Gemma 4 26B A4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 26B A4B IT

Amazon Bedrock

CoverageBenchmark

The Simplismart comparison article explicitly targets Gemma 4 26B A4B (MoE), making it the most technically detailed candidate for this exact variant and the primary source for architecture and deployment characterization. It introduces the model as Google DeepMind's first Mixture-of-Experts entry in the Gemma 4 family On the deployment side, the piece specifies that the 26B A4B can be served on a single NVIDIA A100, details the Official Sampling & Deployment Configuration, and claims a Throughput Advantage over the 31B dense variant owing to the active-parameter design. The article was published August 30, 2026, so it postdates the

Amazon Bedrock

CoverageBenchmark

James Tang benchmarked Gemma 4 26B-A4B on an NVIDIA DGX Spark, explicitly naming the variant as a 4-bit quantized Mixture-of-Experts model with 3.8B active parameters per token. He compared two inference engines: llama.cpp with MXFP4 GGUF and vLLM with NVFP4, both running in single-batch conditions. From the DGX Spark' In measured runs, llama.cpp sustained about 51.57 tokens per second for a single concurrent request, hitting roughly 35% of the theoretical maximum, while vLLM produced about 30 tokens per second under the same conditions. Tang notes the bottleneck is memory bandwidth rather than compute, and frames the post as an unop

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT