Model details
Gemma 3n E4b It
Gemma 3n E4b It represents Google's push to bring capable AI to resource-constrained devices without sacrificing performance. Built on the MatFormer architecture, this model cleverly nests a smaller 2B sub-model within its larger framework, allowing developers to trade off quality against latency depending on the task at hand. The design offloads low-utilization matrices from the accelerator, effectively compressing an 8B-parameter model down to run with a memory footprint matching traditional 4B models. Per-Layer Embedding caching further reduces computational overhead, enabling dynamic memory management that selectively activates parameters based on demand.
Trained across more than 140 languages, this model excels at real-time multimodal processing on mobile and edge hardware—phones, laptops, and tablets benefit from its selective parameter loading. Benchmarks show it performing at 0.75 on HumanEval for code generation, 0.67 on multilingual math reasoning, and 0.65 on broad knowledge tasks. Its flexible nested architecture and open weights make it particularly well-suited for privacy-focused, offline-capable applications where on-device AI delivers meaningful advantages.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- google/gemma-3n-e4b-it
- Release date
- Jun 3, 2025
- Last updated
- Jun 3, 2025
- Knowledge cutoff
- 2024-06
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens
Latest news about Gemma 3n E4b It
No articles yet. Fetch the latest news to show it here.