Currently listed through these providers:
Model details
Qwen3 Next 80B A3B Thinking
Qwen3 Next 80B A3B Thinking belongs to the Qwen3-Next family, a new architecture built around two guiding ideas: scaling total parameters while keeping activated compute low, and pushing context length further without losing efficiency. The architecture combines a hybrid attention mechanism with a highly sparse Mixture-of-Experts structure, along with training-stability optimizations and a multi-token prediction mechanism that speeds up inference. On top of the Qwen3-Next-80B-A3B-Base foundation, an 80-billion-parameter model that activates roughly 3 billion parameters at inference time, the Thinking variant is post-trained for chain-of-thought reasoning, making it well suited to tasks that need deliberate multi-step problem solving rather than shallow pattern matching.
Practically, the model is aimed at developers who want reasoning-class quality at a fraction of the active compute cost, especially when working with long inputs where the Qwen3-Next design is reported to deliver substantially higher throughput above 32K tokens. Its open weights make it attractive for self-hosting and fine-tuning, while the sparse activation pattern keeps serving costs manageable relative to dense 70B-class alternatives. It fits naturally into agentic or analytical pipelines that benefit from explicit thinking traces, structured outputs, and tool use, while still benefiting from the efficiency gains of the underlying Qwen3-Next architecture.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- qwen/qwen3-next-80b-a3b-thinking
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 65,536 tokens