Model details
SmolLM3 3B Base
SmolLM3-3B-Base is a 3 billion-parameter decoder-only transformer from HuggingFaceTB designed to balance compact size with strong reasoning ability. Its architecture uses Grouped Query Attention and NoPE, and it was pretrained on roughly 11.2 trillion tokens drawn from web, code, math, and reasoning data through a staged curriculum. The base release is positioned as a foundation for further pretraining experimentation, with an instruct-tuned sibling recommended for chat or instruction-following use cases.
The model natively supports six languages, including English, French, Spanish, German, Italian, and Portuguese, and ships with a 65,536-token context window by default that can be extended toward 128k tokens through YaRN-based extrapolation. Post-training added about 140 billion reasoning tokens followed by supervised fine-tuning and Anchored Preference Optimization, and the weights are released under an OSI-approved Apache 2.0 license. This combination of small parameter count, multilingual coverage, long-context capability, and open availability makes the base model a practical fit for research teams, fine-tuning projects, and on-device deployments where efficiency matters more than peak scale.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- HuggingFaceTB/SmolLM3-3B-Base
- Release date
- Jun 30, 2025
- Last updated
- Jun 30, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.15
Limits
- Output tokens
- 32,768 tokens
- Context window
- 32,768 tokens
Latest news about SmolLM3 3B Base
No articles yet. Fetch the latest news to show it here.