Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Hugging Face logo

Model details

Gemma 3 4B IT

Gemma 3 4B IT is the instruction-tuned, multimodal entry in Google's Gemma 3 family, designed to pair image and text inputs with text-only outputs in a compact open-weight package. It is distributed on the Hugging Face Hub under the google organization as google/gemma-3-4b-it and is built on the Gemma3ForConditionalGeneration architecture, the same vision-language backbone used across the Gemma 3 multimodal line. The variant is explicitly positioned as Google's multimodal extension of Gemma 3 for tasks such as image captioning and visual question answering, while remaining small enough to run locally or in lightweight serving setups. Open-weight availability makes it attractive for teams that want to fine-tune, distill, or self-host without relying on a closed API.

Practically, the model is aimed at developers who need a general-purpose chat and reasoning assistant that can also see images. It advertises structured outputs and function calling, a long context window that supports extended documents and multi-turn conversations, and a broad multilingual reach spanning over 140 languages. The instruction tuning shapes it for chat-style interaction rather than raw completion, with stronger math, reasoning, and dialogue behavior than earlier Gemma generations. This combination of multimodal input, tool use, and an open license makes Gemma 3 4B IT a flexible base for assistants, document understanding pipelines, and on-device or private-cloud deployments where a compact vision-language model is preferred over a larger frontier system.

Hugging Facegoogle/gemma-3-4b-itgemma

Quick Info

Powered by
Provider
Hugging Face
Model key
google/gemma-3-4b-it
Release date
Mar 12, 2025
Last updated
Mar 12, 2025
Knowledge cutoff
2024-08
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.10

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Gemma 3 4B IT pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemma 3 4B IT

Hugging Face

Coverage

The Google-owned "Gemma 3 Release" collection on Hugging Face is the canonical index for the Gemma 3 model family and explicitly lists the Gemma 3 4B instruction-tuned variant as an Image-Text-to-Text model updated March 21, 2025, alongside its pre-trained sibling, with reported download counts of 1.7M and 88.6k respec The page also surfaces later Google Gemma-family entries — Gemma 4, DiffusionGemma, TranslateGemma, AlphaGenome, Gemma Scope 2, T5Gemma 2, FunctionGemma, EmbeddingGemma, Gemma 3n, MedGemma, Concept Apps, VideoPrism, TxGemma, ShieldGemma, SigLIP2, PaliGemma 2, MetricX-23/24, Gemma 2 JPN, Gemma-APS, TimesFM, and Recurren

Videos about Gemma 3 4B IT

More models around Gemma 3 4B IT