Model details
Qwen3 4B Base
Qwen3-4B-Base is a 4-billion-parameter foundational model from Alibaba's Qwen team, positioned as a customizable starting point for developers rather than a finished assistant. It sits inside the broader Qwen3 lineup, which spans 0.6B to 30B parameters and includes both dense models and Mixture-of-Experts variants, giving builders a consistent family to scale across edge, workstation, and enterprise workloads. Because it is a base checkpoint, the practical fit is for fine-tuning, continued pre-training, or embedding-style adaptation rather than out-of-the-box chat use.
The architecture follows a dense transformer design with 36 layers, which is the same backbone used to derive the separately released Qwen3-Embedding-4B feature-extraction model. The wider Qwen3 family was trained on roughly 36 trillion tokens spanning 119 languages and supports context windows up to 32k tokens, making Qwen3-4B-Base well suited for multilingual experimentation, retrieval pipelines, and domain-specific tuning where long context and broad language coverage matter. Developers get a mid-sized, developer-friendly foundation that can be adapted into specialized chat, reasoning, coding, or retrieval-oriented models while staying inside a single parameter budget.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- Qwen/Qwen3-4B-Base
- Release date
- Mar 31, 2025
- Last updated
- Mar 31, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.15
Limits
- Output tokens
- 32,768 tokens
- Context window
- 32,768 tokens
Latest news about Qwen3 4B Base
No articles yet. Fetch the latest news to show it here.