Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

Gemma 3 12B

Gemma 3 12B is a multimodal language model developed by Google DeepMind, positioned as a well-balanced mid-sized option within the Gemma family. With 12 billion parameters organized across 48 layers and a hybrid attention scheme that pairs full attention layers with sliding window attention, it is designed to handle specialized professional workloads while remaining computationally accessible. The model supports both text and image inputs, producing text outputs, and it can be run locally thanks to its open-weight availability and supported quantization paths that allow deployment on consumer-grade GPUs.

Gemma 3 12B converts visual inputs into tokens and uses an adaptive Pan and Scan approach to preserve detail across images of varying aspect ratios at resolutions up to roughly 896 by 896 pixels. Its expanded the cataloged API limit token context window enables processing of long documents such as legal texts and scientific articles in a single pass, and multilingual support spans more than 140 languages with an enhanced tokenizer inherited from Gemini 2.0. In benchmark tracking, it posts a strong 0.94 score on GSM8k grade-school math problems in zero-shot evaluation, while showing more modest standings in broader leaderboard rankings, reflecting its design as a practical, locally deployable multimodal model rather than a top-tier generalist.

Merge Gatewaygoogle/gemma-3-12b-itgemma

Quick Info

Powered by
Provider
Merge Gateway
Model key
google/gemma-3-12b-it
Release date
Mar 12, 2025
Last updated
Mar 12, 2025
Knowledge cutoff
2024-08
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.29

Limits

Output tokens
8,192 tokens
Context window
128,000 tokens

Transparent token rates

Compare Gemma 3 12B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 3 12B

Merge Gateway

CoverageBenchmark

Google Cloud published a TPU v6e benchmarking study that explicitly targets Gemma 3 12B (alongside Gemma 3 27B) to characterize how the models behave under structurally different inference workloads. The post reports that on decode-heavy generation tasks Gemma 3 27B plateaus at a 4.12x normalized throughput multiplier The blog translates those results into deployment guidance for the Gemma 3 12B specifically: high-concurrency generation workloads should favor the 12B variant or cap the 27B at 64 concurrent requests per replica, while classification and summarization pipelines can safely use the larger model without throughput penalt

Videos about Gemma 3 12B

More models around Gemma 3 12B