Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Gemma 3 4B IT

Gemma 3 4B IT is the instruction-tuned, multimodal entry in Google's Gemma 3 family, designed to pair image and text inputs with text-only outputs in a compact open-weight package. It is distributed on the Hugging Face Hub under the google organization as google/gemma-3-4b-it and is built on the Gemma3ForConditionalGeneration architecture, the same vision-language backbone used across the Gemma 3 multimodal line. The variant is explicitly positioned as Google's multimodal extension of Gemma 3 for tasks such as image captioning and visual question answering, while remaining small enough to run locally or in lightweight serving setups. Open-weight availability makes it attractive for teams that want to fine-tune, distill, or self-host without relying on a closed API.

Practically, the model is aimed at developers who need a general-purpose chat and reasoning assistant that can also see images. It advertises structured outputs and function calling, a long context window that supports extended documents and multi-turn conversations, and a broad multilingual reach spanning over 140 languages. The instruction tuning shapes it for chat-style interaction rather than raw completion, with stronger math, reasoning, and dialogue behavior than earlier Gemma generations. This combination of multimodal input, tool use, and an open license makes Gemma 3 4B IT a flexible base for assistants, document understanding pipelines, and on-device or private-cloud deployments where a compact vision-language model is preferred over a larger frontier system.

OpenRoutergoogle/gemma-3-4b-itgemma

Quick Info

Powered by
Provider
OpenRouter
Model key
google/gemma-3-4b-it
Release date
Mar 12, 2025
Last updated
Mar 12, 2025
Knowledge cutoff
2024-08
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.10

Limits

Output tokens
16,384 tokens
Context window
131,072 tokens

Transparent token rates

Compare Gemma 3 4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 3 4B IT

OpenRouter

CoverageBenchmark

Google: Gemma 3 4B by Google. 131K context, from $0.0400/1M tokens, vision. See benchmarks, comparisons, and our expert ranking — updated 2026.

OpenRouter

Coverage

Physical Intelligence Team Unveils MEM for Robots: A Multi-Scale Memory System Giving Gemma 3-4B VLAs 15-Minute Context for Complex Tasks

Hugging Face

Coverage

The Google-owned "Gemma 3 Release" collection on Hugging Face is the canonical index for the Gemma 3 model family and explicitly lists the Gemma 3 4B instruction-tuned variant as an Image-Text-to-Text model updated March 21, 2025, alongside its pre-trained sibling, with reported download counts of 1.7M and 88.6k respec The page also surfaces later Google Gemma-family entries — Gemma 4, DiffusionGemma, TranslateGemma, AlphaGenome, Gemma Scope 2, T5Gemma 2, FunctionGemma, EmbeddingGemma, Gemma 3n, MedGemma, Concept Apps, VideoPrism, TxGemma, ShieldGemma, SigLIP2, PaliGemma 2, MetricX-23/24, Gemma 2 JPN, Gemma-APS, TimesFM, and Recurren

OpenRouter

Official sourceBenchmark

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. $0.04 per million input tokens, $0.08 per million output tokens. 131,072 token context window, maximum output of 16,384 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.

OpenRouter

Official sourceBenchmark

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. $0 per million input tokens, $0 per million output tokens. 32,768 token context window, maximum output of 8,192 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.

OpenRouter

Coverage

Access ChatGPT, Claude, Gemini & 350+ AI models in one platform. Generate images, videos, music & more. Get started today., Use Google: Gemma 3 4B (free) by Google on Krater.ai. Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles cont... Access 350+ AI models including Google

Videos about Gemma 3 4B IT

More models around Gemma 3 4B IT