Nvidia
Google DeepMind released DiffusionGemma 26B-A4B on June 10, 2026, as an open-weights discrete text diffusion model built on the Gemma 4 26B-A4B Mixture-of-Experts backbone. Rather than predicting one token at a time, it denoises 256-token blocks in parallel across several steps, producing a full block per forward pass for up to 4x faster generation than comparable Gemma 4 autoregressive models on a single NVIDIA H100, reaching more than 1,000 tokens per second. The model has 25.2 billion total parameters with 3.8 billion active per token, a 256K-token context window, multimodal input support covering text, image, and video, more than 140 languages, and a January 2025 knowledge cutoff. It ships under the Apache 2.0 license, allowing commercial and local use, and was created by DeepMind researchers Brendan O'Donoghue and Sebastian Flennerhag who applied the lab's earlier Gemini Diffusion research to the Gemma 4 architecture.