Model details
Qwen3 1.7B Base
As a member of the Qwen3 family, Qwen3 1.7B Base is a causal language model positioned as a pretraining-stage checkpoint rather than a chat-tuned assistant, making it well suited for developers and researchers who want a starting point for domain adaptation, continued training, or experimentation. The broader Qwen3 generation introduces advances over Qwen2.5, including an expanded pretraining corpus of 36 trillion tokens covering 119 languages, with a richer mix of coding, STEM, reasoning, multilingual, and synthetic data. Training is organized into three stages that progress from broad language modeling and general knowledge acquisition, to reasoning-focused skills such as STEM, coding, and logical reasoning, and finally to long-context comprehension with sequence lengths extended up to 32k tokens.
Qwen3 1.7B Base carries the architectural refinements applied across the Qwen3 lineup, including qk layernorm for improved stability, and is configured with 1.7 billion total parameters of which 1.4 billion are non-embedding, arranged across 28 layers. Hyperparameters across the three-stage pipeline were tuned using scaling-law studies, with separate schedules for dense and MoE variants, helping the model achieve consistent dynamics at smaller scales. In practical terms, the 32k token context window combined with a compact parameter footprint makes this base checkpoint a balanced choice for text completion, summarization, and generation workloads where instruction following is not required, while leaving room for targeted fine-tuning in specialized domains.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- Qwen/Qwen3-1.7B-Base
- Release date
- Mar 31, 2025
- Last updated
- Mar 31, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.10
Limits
- Output tokens
- 32,768 tokens
- Context window
- 32,768 tokens
Latest news about Qwen3 1.7B Base
No articles yet. Fetch the latest news to show it here.