Currently listed through these providers:
Model details
Gemma 4 31B IT FP8
Gemma 4 31B IT FP8 is a publicly released instruction-tuned language model built on the latest Gemma architecture and hosted on Hugging Face under the RedHatAI organization. With roughly 31.3 billion parameters, it applies FP8 block quantization to trim memory usage while preserving output quality, allowing it to handle long-form conversations and complex reasoning without truncation. The model's "it" suffix signals optimization for interactive tasks, and third-party commentary describes it as outperforming comparable 31B alternatives on common benchmarks, positioning it as a practical middle-ground choice between smaller open models and much larger frontier systems.
The model accepts both image and text input and produces text output, supporting vision-language workflows alongside traditional chat and reasoning use cases. Strong community traction, reflected in millions of Hugging Face downloads and sustained likes, suggests it has become a go-to open artifact for developers and researchers rather than a niche release. FP8 block quantization makes it possible to run inference on a single high-memory accelerator or a small multi-GPU setup, broadening access for teams that want strong multimodal and reasoning capability without committing to the largest proprietary models.
Quick Info
Powered by- Provider
- InferX
- Model key
- gemma-4-31B-it-fp8
- Release date
- Apr 2, 2026
- Last updated
- Apr 2, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens