Sulat.com
AI models
NaN logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is a multimodal, instruction-tuned mixture-of-experts model that Google has packaged for delivery through NVIDIA NIM, making the weights directly accessible to enterprise inference pipelines on the NGC catalog. The NGC listing frames the architecture as a Mixture-of-Experts design with an instruction-tuned post-training stage, and the available NIM container is shipped as a signed BF16-1.0 build so teams can verify image integrity before deploying. Because the artifact is published under Google's NGC organization path, it slots cleanly into NIM-based serving environments that already standardize on Google's open model lineage.

In practical terms, the combination of multimodal input handling and MoE efficiency makes the model a fit for assistants and document or image-aware workflows that benefit from selective expert routing rather than dense compute at every token. The signed BF16 container and NIM packaging target production scenarios where reproducible deployments and enterprise support matter, including retrieval-augmented chat, structured analysis of mixed text-and-image inputs, and tool-assisted reasoning flows. Teams choosing this variant are typically weighing the trade-off between a sparse-expert footprint and the maturity of the NIM runtime as a serving stack.

NaNgemma4gemma

Quick Info

Powered by
Provider
NaN
Model key
gemma4
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Gemma 4 26B A4B IT

NaN

Coverage

A deployment guide from Flaviu Vladici details how to serve LLMs on the NVIDIA DGX Spark (GB10, sm_121) with vLLM, with Gemma 4 (including the 26B-A4B variant) included as one of three example workloads alongside Qwen and Nemotron. The playbook highlights that the GB10 lacks native FP4 compute, so NVFP4 MoE models must The article also reports measured single-stream throughput on the GB10: ~56 tok/s for Nemotron-Nano, ~52 tok/s for Gemma-26B-A4B, and ~6 tok/s for dense Gemma-31B, plus concrete copy-paste docker run recipes for Gemma 4 (+coder), Qwen3.6, and Nemotron-3 Nano/Omni. It covers speculative decoding two ways (built-in MTP f

NaN

CoverageBenchmark

The OpenRouter model card for google/gemma-4-26b-a4b-it:free documents Gemma 4 26B A4B IT as an instruction-tuned MoE model from Google DeepMind with 25.2B total parameters and 3.8B active per token at inference, targeting near-31B quality at lower compute cost. It supports multimodal input including text, images, and Live provider telemetry on the page shows Google AI Studio delivering a P50 latency of ~0.87s, throughput of 43 tok/s, and 99.51% uptime, with the aggregated routing summary recording 100% uptime and 98.86% availability over the trailing three days. Reported benchmarks include GPQA Diamond scores of 77.2% (SiliconFlow)

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT