Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Gemma 3 12B

Gemma 3 12B belongs to Google's Gemma 3 family of open-weight models, positioned between the smaller Gemma 3 4B sibling and larger Gemma 3 variants. Third-party developer tooling treats it as a vision-capable model that can accept images for OCR, image captioning, and open-ended visual prompts alongside typical text interactions, placing it in the same multimodal evaluation cohort as other mid-size Gemma 3 releases. That positioning suggests a design intent aimed at teams that want on-demand image understanding without committing to a much heavier frontier-scale system, while still benefiting from the broader Gemma lineage's instruction-tuned behavior.

In practical terms, the model fits workflows that mix visual inputs with textual reasoning, such as document extraction, screenshot interpretation, and lightweight image-grounded chat, where its multimodal grounding is the main differentiator over purely text-only models of similar size. Open weights allow self-hosting, fine-tuning, and auditing, which suits regulated or cost-sensitive deployments that prefer to run a known recipe locally rather than rely solely on a hosted endpoint. For developers comparing vision-capable open-weight options, Gemma 3 12B offers a middle ground between the lighter Gemma 3 4B for low-latency tasks and larger closed or open models reserved for the most demanding multimodal reasoning workloads.

NovitaAIgoogle/gemma-3-12b-itgemma

Quick Info

Powered by
Provider
NovitaAI
Model key
google/gemma-3-12b-it
Release date
Mar 13, 2025
Last updated
Mar 13, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.10

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Gemma 3 12B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 3 12B

NovitaAI

CoverageBenchmark

Google Cloud published a TPU v6e benchmarking study on September 4, 2026 explicitly comparing Gemma 3 12B against Gemma 3 27B across structurally distinct LLM workloads. The study found that for decode-heavy generation tasks, the Gemma 3 12B scales up to an 8.19x normalized throughput multiplier at 128 concurrent users For prefill-heavy classification workloads, the Google Cloud post reports that model parameter size matters far less — both Gemma 3 12B and 27B achieve similar peak scaling of roughly 6.0x to 6.4x normalized throughput at 128 users without saturating the TPUs, suggesting larger models can be deployed for classification

Videos about Gemma 3 12B

More models around Gemma 3 12B