Currently listed through these providers:
Model details
Gemma 4 E2B IT
Gemma 4 E2B IT is positioned as the most resource-efficient member of the Gemma 4 family, engineered to deliver capable reasoning and multimodal understanding under tight hardware budgets. The model pairs a 5.1-billion-parameter total design with roughly 2.3 billion parameters active during inference, using Per-Layer Embeddings to keep memory pressure low while preserving quality. It is structured around 35 layers and employs a hybrid attention scheme with a 512-token sliding window, supporting a 128,000-token context window for sustained dialogue and document tasks. Its multimodal reach covers text and images natively, with audio handled through a dedicated encoder of approximately 300 million parameters, making it suitable for assistants that need to interpret spoken input alongside visual and written cues.
For practitioners, the practical draw of Gemma 4 E2B IT is its ability to bring Gemma-class reasoning onto laptops, phones, and other constrained devices without sacrificing multimodal breadth. Independent on-device benchmarks on Qualcomm's Snapdragon X2 Elite and Snapdragon X Elite platforms, using a q4_0 quantized llama.cpp distribution, report prefilling speeds around 1,808 tokens per second, decoding near 35.1 tokens per second, and an on-device accuracy of 93.3% relative to a 95.8% reference — a credible fidelity gap for a heavily compressed variant. Developers highlight agentic workflows, OCR, and latency-sensitive scenarios as especially well-matched use cases, and the Apache 2.0 license broadens options for commercial integration. Together, these traits position E2B as a pragmatic bridge between full-scale Gemma reasoning and the realities of edge deployment.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- google.gemma-4-e2b
- Release date
- Apr 2, 2026
- Last updated
- Apr 2, 2026
- AI SDK package
@ai-sdk/amazon-bedrock/mantle- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.04
- Output token cost
- $0.08
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Gemma 4 E2B IT pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemma 4 E2B IT
No articles yet. Fetch the latest news to show it here.