Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

Gemma 4 E4B IT

Gemma 4 E4B IT belongs to the Gemma 4 family of open models built by Google DeepMind and released under the Apache 2.0 license, joining a lineup that includes five sizes ranging from the compact E2B up to a 31B variant. The family mixes Dense and Mixture-of-Experts architectures so the same weights can serve use cases from on-device assistants to server-side inference, and the open-weights approach is reinforced with instruction-tuned variants like this one for chat and tool-driven applications. Beyond standard text, the model accepts image input and adds native audio understanding on the E2B, E4B, and 12B variants, making it useful for assistants that need to listen, read, and see within a single response.

The instruction-tuned E4B variant is positioned as a capable generalist reasoner with configurable thinking modes, an extended context window of up to 256K tokens, and multilingual support across more than 140 languages, which broadens its fit for retrieval-heavy assistants, long-document analysis, and code generation. Deep Infra exposes the model as a server-side text generation endpoint with priority and flex throughput tiers, so teams can choose between faster interactive responses and lower-cost batch-style workloads. Combined with its multimodal inputs and reasoning orientation, Gemma 4 E4B IT is well matched to production assistants, developer tools, and multilingual chatbots where open weights and flexible deployment matter.

Deep Infragoogle/gemma-4-E4B-itgemmadeprecated

Quick Info

Powered by
Provider
Deep Infra
Model key
google/gemma-4-E4B-it
Release date
Apr 2, 2026
Last updated
Apr 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.02
Output token cost
$0.10

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Gemma 4 E4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 4 E4B IT

Pioneer

Coverage

The official Google DeepMind model card on Hugging Face for google/gemma-4-E4B-it describes Gemma 4 as a family of open-weight models released under the Apache 2.0 license, with the E4B variant being one of five distinct sizes (E2B, E4B, 12B, 26B A4B, and 31B). The E4B-it card confirms multimodal text-and-image input, The card situates E4B-it between the phone-sized E2B and the workstation-class 12B, 26B-A4B (MoE), and 31B (dense) variants, targeting laptops and high-end phones as its primary deployment environments. It emphasizes hybrid attention that interleaves local sliding-window attention with full global attention, Per-Layer

Videos about Gemma 4 E4B IT

More models around Gemma 4 E4B IT