Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3 8B

Qwen3 8B is a dense causal language model built with 8.2 billion parameters and a 36-layer architecture. It is designed to provide a flexible user experience by allowing seamless switching between a specialized thinking mode for complex logical, mathematical, and coding tasks and a standard non-thinking mode for general-purpose dialogue. This dual-mode capability ensures the model remains efficient for everyday interactions while maintaining the depth required for intricate problem-solving and agent-based workflows, where it can integrate precisely with external tools.

The model is grounded in a robust training foundation that utilized approximately 36 trillion tokens of high-quality multilingual data, including web text, technical documentation, and synthetic domain-specific content. Following its initial pre-training, the model underwent a rigorous four-stage reinforcement process to refine its instruction-following and human preference alignment. These advancements enable the model to excel in creative writing, role-playing, and translation across more than 100 languages. Its compact size and strong performance make it a practical choice for on-device deployment and intelligent applications that require reliable reasoning and agentic intelligence.

Alibabaqwen3-8bqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3-8b
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.18
Output token cost
$0.70

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3 8B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 8B

Alibaba

CoverageBenchmark

A July 25, 2026 Medium architecture and deployment guide by Basanta Sapkota explicitly addresses Qwen3-8B-Base, the open-weight dense variant that is the same family and parameter class as the Qwen3 8B subject. The article states Qwen3-8B-Base contains 8.2 billion total parameters with 6.95 billion non-embedding parame The same source adds Qwen3-8B pretraining specifics relevant to the subject: the model was pre-trained on 36 trillion tokens spanning 119 languages using a three-stage pipeline that extended training sequences to 32,768 tokens, and all Qwen3 models incorporate QK LayerNorm (with MoE variants adding global-batch load ba

Videos about Qwen3 8B

More models around Qwen3 8B