Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Qwen3 8B

Qwen3-8B belongs to the latest Qwen3 generation of large language models, a suite that includes both dense and mixture-of-experts variants. As a dense 8.19 billion parameter model with about 6.95 billion non-embedding parameters across 36 layers, it sits in a middle-weight tier designed to balance capability with efficiency. The model was released with open weights under the Apache License 2.0, making it available for local deployment, fine-tuning, and redistribution. Its architecture uses grouped-query attention with 32 query heads and 8 key-value heads, a configuration that helps manage memory and inference cost at this scale.

A defining feature of Qwen3-8B is its seamless switching between a thinking mode, intended for complex logical reasoning, mathematics, and coding, and a non-thinking mode optimized for efficient general-purpose dialogue. The upstream Qwen team highlights that this reasoning-focused mode surpasses earlier QwQ models on math, code generation, and commonsense reasoning, while the non-thinking mode improves on prior Qwen2.5 instruct behavior. The model also emphasizes instruction-following, agent capabilities with external tool integration, and broad multilingual coverage spanning over 100 languages. These qualities make Qwen3-8B a practical fit for developers who need a flexible open-weight model that can shift between deliberate reasoning and fast conversational response within a single deployment.

NovitaAIqwen/qwen3-8b-fp8

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen3-8b-fp8
Release date
Apr 29, 2025
Last updated
Apr 29, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.035
Output token cost
$0.138

Limits

Output tokens
20,000 tokens
Context window
128,000 tokens

Latest news about Qwen3 8B

Videos about Qwen3 8B