Sulat.com
AI models
Cloudflare Workers AI logo

Model details

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is a Google open-weight model that uses a mixture-of-experts design with 26 billion total parameters while activating roughly 4 billion per forward pass, giving it a favorable balance between capacity and inference cost. It is described as built on the same underlying architecture that powers Gemini 3, bringing modern alignment and multimodal capabilities to the open Gemma family. The instruction-tuned variant is aimed at assistant-style usage where developers need chat, reasoning, and image understanding in a single deployable checkpoint.

For practical deployment, the model supports text and image inputs, function-calling, and structured JSON output, which makes it useful for building agents, document understanding pipelines, and multilingual applications across 140 or more languages within a long context window. Its sparse activation pattern keeps per-request compute closer to a small model even though the full parameter count is large, which is helpful for throughput-sensitive workloads running on serverless inference. Teams looking for an open-weight alternative to closed flagship assistants, especially those needing vision and tool use without a heavy inference footprint, will find this variant a sensible fit.

Cloudflare Workers AI@cf/google/gemma-4-26b-a4b-itgemma

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/google/gemma-4-26b-a4b-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
16,384 tokens
Context window
256,000 tokens

Latest news about Gemma 4 26B A4B IT

Videos about Gemma 4 26B A4B IT

More models around Gemma 4 26B A4B IT