Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

DiffusionGemma 26B A4B IT

The model overview is being prepared.

Nvidiagoogle/diffusiongemma-26b-a4b-itgemma

Quick Info

Powered by
Provider
Nvidia
Model key
google/diffusiongemma-26b-a4b-it
Release date
Jun 9, 2026
Last updated
Jul 15, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
250,000 tokens

Latest news about DiffusionGemma 26B A4B IT

Nvidia

CoverageRelease Notes

Google DeepMind released DiffusionGemma 26B-A4B on June 10, 2026, as an open-weights discrete text diffusion model built on the Gemma 4 26B-A4B Mixture-of-Experts backbone. Rather than predicting one token at a time, it denoises 256-token blocks in parallel across several steps, producing a full block per forward pass for up to 4x faster generation than comparable Gemma 4 autoregressive models on a single NVIDIA H100, reaching more than 1,000 tokens per second. The model has 25.2 billion total parameters with 3.8 billion active per token, a 256K-token context window, multimodal input support covering text, image, and video, more than 140 languages, and a January 2025 knowledge cutoff. It ships under the Apache 2.0 license, allowing commercial and local use, and was created by DeepMind researchers Brendan O'Donoghue and Sebastian Flennerhag who applied the lab's earlier Gemini Diffusion research to the Gemma 4 architecture.

Nvidia

Coverage

Google released DiffusionGemma as an open model on June 10, 2026, applying a diffusion-based generation process to language models in place of standard autoregressive token-by-token output. Built on results from the earlier Gemini Diffusion project, it uses a Mixture-of-Experts design with 25.2 billion total parameters and 3.8 billion active parameters, delivering what Google describes as overwhelmingly faster generation than Gemma 4 31B, 26B A4B, and 12B with MTP acceleration. Coverage notes that diffusion language models iterate over entire noisy canvases to refine them into final output, enabling high-speed responses suited to latency-sensitive use cases. The article also points to Google's first-party developer guide and blog posts for benchmark data and to a Sudoku-focused fine-tuned variant of the release, while the English text is partly machine-translated from the original Japanese report.

Videos about DiffusionGemma 26B A4B IT

More models around DiffusionGemma 26B A4B IT