Currently listed through these providers:
Model details
Qwen3 8B
Qwen3 8B is a dense causal language model in the Qwen3 series from Alibaba Cloud, designed to switch fluidly between thinking and non-thinking modes within a single architecture. In thinking mode it handles complex logical reasoning, mathematics, and coding tasks, while non-thinking mode delivers efficient general-purpose dialogue. The model achieves these capabilities through extensive pretraining and post-training across 36 layers with grouped-query attention, supporting native 32K context that expands significantly with YaRN. It includes agentic capabilities for precise tool integration in both modes, and covers over 100 languages and dialects for multilingual instruction following and translation.
Training evidence from the sources shows Qwen3 8B underwent pretraining and post-training, with benchmark results surpassing its predecessors QwQ and Qwen2.5 instruct models in mathematics, code generation, and commonsense logical reasoning. The model demonstrates strong human preference alignment for creative writing, role-playing, and multi-turn conversations. Released under the Apache 2.0 license with publicly available technical reports, it offers accessible formats like GGUF and Ollama. With 8.2B total parameters and 6.95B non-embedding parameters, it balances capability and efficiency for open-source deployment across research and production use cases.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- qwen3-8b
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.072
- Output token cost
- $0.287
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3 8B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 8B
No articles yet. Fetch the latest news to show it here.