Currently listed through these providers:
Model details
Gemma 3 27B
Gemma 3 27B is Google's open flagship model in the Gemma family, built from the same research lineage that produced the Gemini models. This instruction-tuned variant carries 27 billion parameters and pairs a language backbone with a SigLIP vision encoder to handle both text and image inputs while generating text output. The model supports a generous 128K context window—among the largest available in open-weight models of this size—and multilingual support across more than 140 languages, making it broadly applicable across regions and use cases. Its architecture balances capability with deployability: the 27B size fits on a single high-end GPU like an H100 in BF16 precision without needing quantization, opening the door for independent researchers, smaller teams, or cost-conscious deployments to run a genuinely competitive model on their own infrastructure.
The model's training leveraged Google's scaled research foundation, and at launch Gemma 3 27B placed in the global Chatbot Arena top 10, outranking open models with far larger parameter counts including Llama 3 405B and DeepSeek V3 671B—achieving this while running on a single GPU. Beyond raw benchmarks, Google reports improvements in math, reasoning, conversational interaction, structured outputs, and function calling compared to its predecessors. These capabilities make it well-suited for production tasks like question answering, summarization, document understanding with embedded images, and multilingual customer service. Deployment options include vLLM and Ollama, with flexibility for Ubuntu-based GPU servers and integration with interfaces like Open Web UI, giving operators full control over data and infrastructure rather than relying on hosted endpoints.
Quick Info
Powered by- Provider
- STACKIT
- Model key
- google/gemma-3-27b-it
- Release date
- May 17, 2025
- Last updated
- May 17, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.53
- Output token cost
- $0.76
Limits
- Output tokens
- 4,096 tokens
- Context window
- 37,000 tokens