Model details
Qwen3-Next-80B-A3B-Thinking-fast
The Qwen3-Next-80B-A3B-Thinking model leverages a sparse mixture-of-experts architecture, routing tasks through 80 billion total parameters with only 3.9 billion active parameters engaged per token. This design choice means the model achieves its capabilities without proportional computational overhead, making dense-model-level performance more accessible for production deployments. The model sits within Alibaba's Qwen3 series of open-weight language models, carrying forward a lineage built on extensive pre-training at scale and iterative refinement across generations.
As an open-weights model, Qwen3-Next-80B-A3B-Thinking exposes its weights publicly for local inference, self-hosting, and fine-tuning, which has driven adoption among developers who need transparency or data privacy guarantees. The architecture is optimized for structured reasoning tasks and tool-augmented workflows, reflecting an intentional shift toward models that can follow multi-step plans and call external functions reliably. Its broad API compatibility allows developers to swap in this model with minimal code changes, effectively lowering the barrier to experimenting with an open-weight frontier-class model in existing pipelines.
Quick Info
Powered by- Provider
- Nebius Token Factory
- Model key
- Qwen/Qwen3-Next-80B-A3B-Thinking-fast
- Release date
- Jul 25, 2025
- Last updated
- May 7, 2026
- Knowledge cutoff
- 2025-07
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.20
Limits
- Input tokens
- 7,000 tokens
- Output tokens
- 8,192 tokens
- Context window
- 8,000 tokens