Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

Qwen/Qwen3-8B

Qwen3-8B is a dense 8.2 billion parameter causal language model built in the Qwen3 series by Alibaba, featuring 36 transformer layers with grouped query attention (32 heads for queries and 8 for key-value pairs). What makes it particularly versatile is its unique ability to switch between a thinking mode for complex logical reasoning, mathematics, and code generation, and a non-thinking mode for efficient general-purpose dialogue — all within a single model. This dual-mode design lets developers and end users choose the right behavior for each scenario without managing separate deployments.

The model was developed through pretraining followed by post-training phases, and its post-training work is credited with delivering reasoning performance that surpasses both its predecessor QwQ and the Qwen2.5 instruct family on mathematics, code generation, and commonsense reasoning benchmarks. Qwen3-8B also stands out for human preference alignment, excelling in creative writing, role-playing, and multi-turn conversations, while its agent capabilities allow it to integrate with external tools in both modes. Supporting over 100 languages and dialects, it targets developers building multilingual, agent-driven, or reasoning-intensive applications — from interactive assistants to complex automation pipelines.

SiliconFlowQwen/Qwen3-8Bqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3-8B
Release date
Apr 30, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.06

Limits

Output tokens
131,000 tokens
Context window
131,000 tokens

Transparent token rates

Compare Qwen/Qwen3-8B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-8B

SiliconFlow (China)

CoverageBenchmark

Reproducible benchmark results for Qwen/Qwen3-8B were evaluated across 13 tasks on a 16.4 GB safetensors checkpoint, providing independent verification of the model's capabilities. The strongest showing was GSM8K at 84.9% exact match (rank 10 of 28), followed by IFEval at 81.7% prompt-level strict accuracy and MMLU at 73.0%. These measurements give developers third-party validation of Qwen3-8B's math, instruction-following, and general knowledge performance. Additional reproducible scores include MBPP at 65.6% pass@1, MMLU-Pro at 57.7%, MGSM at 47.9%, and EQ-Bench at 75.8, while GPQA Diamond reached 27.8% and the MBPP (Instruct) variant scored 0.0%. The model ranked in the top three on none of the 13 benchmarks evaluated, with its weakest results on instruction-formatted MBPP, giving a clear picture of where Qwen3-8B excels versus where it struggles in real coding workloads.

SiliconFlow (China)

CoverageBenchmark

Qwen3-8B is an 8B-parameter dense reasoning model from Alibaba, released on April 28, 2025, and distributed under the Apache 2.0 license with open weights. The cataloged specifications include a 128K-token context window and a 126K-token maximum output, supporting text input and output with reasoning and multilingual capabilities. This third-party catalog independently records the model as deprecated with a scheduled retirement date of October 10, 2026, which is timely signal for users of the SiliconFlow-hosted checkpoint. The same source documents Qwen3-8B benchmark performance across academic, coding, and reasoning categories. Notable recorded scores include AIME 2024 at 79.4, GPQA Diamond at 62.0, MATH-500 at 97.4 under thinking mode, and MMLU-Pro at 72.5. These figures establish the model's relative standing among 8B-class reasoning models and reinforce the case for migration before the listed October 2026 retirement window.

Videos about Qwen/Qwen3-8B

More models around Qwen/Qwen3-8B