Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

Google Gemma 4 26B A4B Instruct

Google Gemma 4 26B A4B Instruct is the flagship server-class member of the Gemma 4 family, designed as a Mixture-of-Experts vision-language model that pairs a roughly 25B-parameter total footprint with only about 3.8B activated parameters per token. This sparse activation pattern is intended to deliver near-4B-class inference latency while preserving the reasoning quality of much larger dense systems, and the architecture supports multimodal understanding across text, image, and video inputs with text outputs. the cataloged API limit context window lets it handle long documents, extended dialogues, and multi-step reasoning chains that would overflow smaller windows, and the model carries explicit reasoning-oriented design choices suitable for analytical and tool-augmented workflows.

As an open-weights release in the Gemma lineage, the 26B A4B Instruct variant is well suited to organizations that want frontier-class reasoning and multimodal comprehension without paying the inference cost of a fully dense large model. It is validated for day-zero performance on Intel Xeon CPUs (generations 4 through 6) running Red Hat, where Intel's upstream FusedMoE kernels in vLLM keep the expert-routing layers efficient out of the box, making it a practical choice for on-premise and hybrid deployments. Teams building document analysis, video understanding, code reasoning, or agentic pipelines benefit from its large context, sparse activation efficiency, and open availability for self-hosting and customization.

Venice AIgoogle-gemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
Venice AI
Model key
google-gemma-4-26b-a4b-it
Release date
Apr 2, 2026
Last updated
Jun 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.13
Output token cost
$0.40

Limits

Output tokens
8,192 tokens
Context window
256,000 tokens

Transparent token rates

Compare Google Gemma 4 26B A4B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Google Gemma 4 26B A4B Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Google Gemma 4 26B A4B Instruct

More models around Google Gemma 4 26B A4B Instruct