Currently listed through these providers:
Model details
Gemma 4 26B-A4B
Gemma 4 26B-A4B is a member of the Gemma 4 family of open-weight models developed by Google DeepMind, designed to balance capable reasoning with efficient serving. It uses a Mixture-of-Experts architecture rather than a purely dense design, sitting within a lineup that also includes dense and MoE variants ranging from compact E2B and E4B sizes up through 12B and 31B options. The 26B-A4B naming indicates an active-parameter configuration that aims to deliver stronger reasoning performance while keeping inference costs lower than a full 26B dense model would require.
The model is multimodal on the input side, accepting text and image, and producing text output, making it well suited to tasks like document understanding, visual question answering, code generation, and general reasoning workloads. It supports a long context window of up to 256K tokens and maintains multilingual coverage across more than 140 languages, which helps with translation, cross-lingual reasoning, and long-document analysis. Released under the Apache 2.0 license in pre-trained and instruction-tuned forms, Gemma 4 26B-A4B fits practitioners who want an open-weight Google DeepMind model that can be self-hosted or fine-tuned for specialized assistants, while still offering the multimodal grounding and extended context expected from a modern reasoning model.
Quick Info
Powered by- Provider
- EmpirioLabs AI
- Model key
- gemma-4-26b-a4b
- Release date
- Apr 2, 2026
- Last updated
- Jun 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.29
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare gemma pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemma 4 26B-A4B
No articles yet. Fetch the latest news to show it here.