Currently listed through these providers:
Model details
Google Gemma 4 26B A4B Instruct
Google Gemma 4 26B A4B Instruct is the flagship server-class member of the Gemma 4 family, designed as a Mixture-of-Experts vision-language model that pairs a roughly 25B-parameter total footprint with only about 3.8B activated parameters per token. This sparse activation pattern is intended to deliver near-4B-class inference latency while preserving the reasoning quality of much larger dense systems, and the architecture supports multimodal understanding across text, image, and video inputs with text outputs. the cataloged API limit context window lets it handle long documents, extended dialogues, and multi-step reasoning chains that would overflow smaller windows, and the model carries explicit reasoning-oriented design choices suitable for analytical and tool-augmented workflows.
As an open-weights release in the Gemma lineage, the 26B A4B Instruct variant is well suited to organizations that want frontier-class reasoning and multimodal comprehension without paying the inference cost of a fully dense large model. It is validated for day-zero performance on Intel Xeon CPUs (generations 4 through 6) running Red Hat, where Intel's upstream FusedMoE kernels in vLLM keep the expert-routing layers efficient out of the box, making it a practical choice for on-premise and hybrid deployments. Teams building document analysis, video understanding, code reasoning, or agentic pipelines benefit from its large context, sparse activation efficiency, and open availability for self-hosting and customization.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- google-gemma-4-26b-a4b-it
- Release date
- Apr 2, 2026
- Last updated
- Jun 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.13
- Output token cost
- $0.40
Limits
- Output tokens
- 8,192 tokens
- Context window
- 256,000 tokens
Transparent token rates
Compare Google Gemma 4 26B A4B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Google Gemma 4 26B A4B Instruct
No articles yet. Fetch the latest news to show it here.