Model details
Qwen3 8B
Qwen3-8B belongs to the latest Qwen3 generation of large language models, a suite that includes both dense and mixture-of-experts variants. As a dense 8.19 billion parameter model with about 6.95 billion non-embedding parameters across 36 layers, it sits in a middle-weight tier designed to balance capability with efficiency. The model was released with open weights under the Apache License 2.0, making it available for local deployment, fine-tuning, and redistribution. Its architecture uses grouped-query attention with 32 query heads and 8 key-value heads, a configuration that helps manage memory and inference cost at this scale.
A defining feature of Qwen3-8B is its seamless switching between a thinking mode, intended for complex logical reasoning, mathematics, and coding, and a non-thinking mode optimized for efficient general-purpose dialogue. The upstream Qwen team highlights that this reasoning-focused mode surpasses earlier QwQ models on math, code generation, and commonsense reasoning, while the non-thinking mode improves on prior Qwen2.5 instruct behavior. The model also emphasizes instruction-following, agent capabilities with external tool integration, and broad multilingual coverage spanning over 100 languages. These qualities make Qwen3-8B a practical fit for developers who need a flexible open-weight model that can shift between deliberate reasoning and fast conversational response within a single deployment.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- qwen/qwen3-8b-fp8
- Release date
- Apr 29, 2025
- Last updated
- Apr 29, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.035
- Output token cost
- $0.138
Limits
- Output tokens
- 20,000 tokens
- Context window
- 128,000 tokens