Gemma 3 27B is Google's flagship open-weight entry in the Gemma 3 family, a multimodal vision-language model that accepts both text and image input while producing text output. It belongs to a lineup of five instruction-tuned variants ranging from 270M to 27B parameters, with the 4B, 12B, and 27B versions sharing multimodal capability and a 128K-token context window. The 27B variant in BF16 occupies roughly 54 GB of VRAM including its SigLIP vision encoder, making it feasible to run on a single high-end GPU such as an H100 without quantization, though production workloads benefit from around 80 GB of headroom to accommodate KV cache and longer contexts.
At launch, Gemma 3 27B ranked in the global Chatbot Arena top ten and was reported to outperform considerably larger open models, positioning it as a strong choice for high-quality production deployments that demand both language understanding and visual reasoning. Beyond raw capability, the model supports over 140 languages and offers structured outputs and function calling, broadening its applicability across diverse enterprise use cases. Its combination of open weights, multimodal input, a long context window, and single-GPU footprint makes it well suited for teams seeking a capable general-purpose model without the infrastructure overhead of much larger frontier systems.