Currently listed through these providers:
Model details
Gemma 4 26B A4B IT
Gemma 4 26B A4B IT is a Google open-weight model that uses a mixture-of-experts design with 26 billion total parameters while activating roughly 4 billion per forward pass, giving it a favorable balance between capacity and inference cost. It is described as built on the same underlying architecture that powers Gemini 3, bringing modern alignment and multimodal capabilities to the open Gemma family. The instruction-tuned variant is aimed at assistant-style usage where developers need chat, reasoning, and image understanding in a single deployable checkpoint.
For practical deployment, the model supports text and image inputs, function-calling, and structured JSON output, which makes it useful for building agents, document understanding pipelines, and multilingual applications across 140 or more languages within a long context window. Its sparse activation pattern keeps per-request compute closer to a small model even though the full parameter count is large, which is helpful for throughput-sensitive workloads running on serverless inference. Teams looking for an open-weight alternative to closed flagship assistants, especially those needing vision and tool use without a heavy inference footprint, will find this variant a sensible fit.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/google/gemma-4-26b-a4b-it
- Release date
- Apr 2, 2026
- Last updated
- Apr 2, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.30
Limits
- Output tokens
- 16,384 tokens
- Context window
- 256,000 tokens
Latest news about Gemma 4 26B A4B IT
Videos about Gemma 4 26B A4B IT
More models around Gemma 4 26B A4B IT
This exact model name is also listed by 13 other providers.