Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pioneer logo

Model details

DiffusionGemma 26B-A4B IT

DiffusionGemma 26B-A4B IT is positioned as a sizable language model built around an active-parameter design hinted at by the A4B suffix, with a third-party infrastructure page estimating roughly 25.8 billion parameters. That same page derives practical guidance for running it, placing full-precision inference needs in the neighborhood of 56 GB of VRAM at a 4k context window and batch size of one, while quantized deployments can fit on smaller GPU configurations. The model also appears in NVIDIA's NGC catalog as an NIM container under the google team path, suggesting it is packaged for deployment on optimized inference runtimes alongside other recent open and proprietary releases.

In practical terms, the model is aimed at reasoning-heavy and tool-augmented tasks where temperature control and structured tool calling matter, and the large context window supports extended multi-turn conversations and long-form analysis. Its roughly 26 billion active parameter count sits in a sweet spot between lighter assistants and the largest frontier models, making it a reasonable choice for teams that want strong reasoning behavior without committing to top-tier hardware. Developers considering self-hosting can plan around the GPU recommendations from Spheron, balancing FP16 fidelity against INT8 or INT4 quantization to match available accelerators and budget constraints.

Pioneergoogle/diffusiongemma-26B-A4B-it

Quick Info

Powered by
Provider
Pioneer
Model key
google/diffusiongemma-26B-A4B-it
Release date
May 31, 2026
Last updated
May 31, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Latest news about DiffusionGemma 26B-A4B IT

Pioneer

CoverageRelease Notes

MLQ News reports on the June 10, 2026 Google DeepMind release of DiffusionGemma, a 26B open-weights model that generates text via discrete diffusion — the same technique underpinning image generators like Stable Diffusion — producing up to 256 tokens in parallel per forward pass. The article confirms Apache 2.0 distrib The piece documents that DiffusionGemma achieves over 1,000 tokens per second on a single NVIDIA H100, roughly four to five times faster than comparable autoregressive models. Output quality is reported as lower than standard Gemma 4 on benchmarks such as MMLU and coding tests, with Google positioning the model as expe

Pioneer

Coverage

Google released DiffusionGemma as an open model on June 10, 2026, positioning it as a diffusion-based language model built on the Gemini Diffusion research line. Unlike autoregressive models that emit tokens sequentially, DiffusionGemma iteratively refines outputs across the full sequence, a paradigm translated from im The article details concrete architecture: a Mixture-of-Experts model with roughly 25.2 billion total parameters and 3.8 billion active parameters (the "26B-A4B" configuration). In benchmark comparisons against Gemma 4 31B, Gemma 4 26B A4B, and Gemma 4 12B, DiffusionGemma reportedly produces dramatically more output to

Pioneer

CoverageRelease Notes

Google AI, including Google DeepMind researchers, released DiffusionGemma on June 10, 2026 as an experimental open-weights model that uses text diffusion rather than standard autoregressive decoding for generation. The model is shipped under the Apache 2.0 license, making it available to developers and researchers expl MarkTechPost confirms DiffusionGemma's architecture as a 26B-parameter Mixture-of-Experts text-diffusion model and highlights Google's claim of up to 4x faster generation compared with conventional autoregressive approaches. The article frames the release as part of an emerging line of diffusion language models and pos

Pioneer

Coverage

Google DeepMind released DiffusionGemma, a 26B-parameter Mixture-of-Experts open-weights model released under Apache 2.0, on June 10, 2026. The model uses discrete text diffusion to generate 256-token blocks in parallel per forward pass, shifting away from token-by-token autoregression. With 3.8B active parameters duri Google's headline performance claims include 1,000+ tokens per second on a single NVIDIA H100 and 700+ tokens per second on an RTX 5090, described as roughly four to five times faster than comparable autoregressive models. The source explicitly notes these speedups are designed for local and low-concurrency inference,

Pioneer

Coverage

The ModelScope listing for google/diffusiongemma-26B-A4B-it carries Google DeepMind authorship and serves as the first-party model card mirror under Apache 2.0 license, with 25.82B parameters in the Transformers/Safetensors/PyTorch format. DiffusionGemma is described as a generative model built by Google DeepMind based Key technical capabilities explicitly documented include configurable Thinking Mode for reasoning, native system prompt support inherited from Gemma 4, and specific optimization for small-batch, low-latency, single-accelerator inference. The listing indicates the model is positioned to reduce the sequential bottlenecks

Videos about DiffusionGemma 26B-A4B IT