Sulat.com
AI models
InferX logo

Model details

Gemma 4 31B IT FP8

Gemma 4 31B IT FP8 is a publicly released instruction-tuned language model built on the latest Gemma architecture and hosted on Hugging Face under the RedHatAI organization. With roughly 31.3 billion parameters, it applies FP8 block quantization to trim memory usage while preserving output quality, allowing it to handle long-form conversations and complex reasoning without truncation. The model's "it" suffix signals optimization for interactive tasks, and third-party commentary describes it as outperforming comparable 31B alternatives on common benchmarks, positioning it as a practical middle-ground choice between smaller open models and much larger frontier systems.

The model accepts both image and text input and produces text output, supporting vision-language workflows alongside traditional chat and reasoning use cases. Strong community traction, reflected in millions of Hugging Face downloads and sustained likes, suggests it has become a go-to open artifact for developers and researchers rather than a niche release. FP8 block quantization makes it possible to run inference on a single high-memory accelerator or a small multi-GPU setup, broadening access for teams that want strong multimodal and reasoning capability without committing to the largest proprietary models.

InferXgemma-4-31B-it-fp8gemma

Quick Info

Powered by
Provider
InferX
Model key
gemma-4-31B-it-fp8
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Gemma 4 31B IT FP8

Videos about Gemma 4 31B IT FP8

More models around Gemma 4 31B IT FP8