Gemma 3 12B is a multimodal language model developed by Google DeepMind, positioned as a well-balanced mid-sized option within the Gemma family. With 12 billion parameters organized across 48 layers and a hybrid attention scheme that pairs full attention layers with sliding window attention, it is designed to handle specialized professional workloads while remaining computationally accessible. The model supports both text and image inputs, producing text outputs, and it can be run locally thanks to its open-weight availability and supported quantization paths that allow deployment on consumer-grade GPUs.
Gemma 3 12B converts visual inputs into tokens and uses an adaptive Pan and Scan approach to preserve detail across images of varying aspect ratios at resolutions up to roughly 896 by 896 pixels. Its expanded the cataloged API limit token context window enables processing of long documents such as legal texts and scientific articles in a single pass, and multilingual support spans more than 140 languages with an enhanced tokenizer inherited from Gemini 2.0. In benchmark tracking, it posts a strong 0.94 score on GSM8k grade-school math problems in zero-shot evaluation, while showing more modest standings in broader leaderboard rankings, reflecting its design as a practical, locally deployable multimodal model rather than a top-tier generalist.