Currently listed through these providers:
Model details
Qwen3 Coder Next FP8
Qwen3 Coder Next FP8 is an 80-billion parameter Mixture-of-Experts model that activates only 3 billion parameters during inference, drawing from a pool of 512 experts with 10 active plus one shared expert per forward pass. Its architecture combines Gated DeltaNet linear attention with Gated Attention layers across 48 transformer layers, creating a hybrid design purpose-built for coding workflows. The 256K-token context window supports long files and multi-file projects, while fine-grained FP8 quantization with a block size of 128 enables efficient deployment on accessible hardware. This non-thinking mode model is engineered specifically for coding agents and local development scenarios.
The model underwent both pretraining and post-training stages through an elaborate recipe designed to cultivate coding capabilities. It achieves performance on par with models requiring 10 to 20 times more activated parameters, making it a cost-effective choice for agent-based deployments. The training approach emphasizes long-horizon reasoning, complex tool usage, and recovery from execution failures—competencies that show up directly in real-world coding tasks. Its adaptability to various scaffold templates enables seamless integration with popular CLI and IDE platforms, supporting diverse development environments from the ground up. Released under the Apache 2.0 license, this open-weight checkpoint brings the efficiency of MoE architecture to coding agents and individual developers alike.
Quick Info
Powered by- Provider
- Together AI
- Model key
- Qwen/Qwen3-Coder-Next-FP8
- Release date
- Feb 3, 2026
- Last updated
- Feb 3, 2026
- Knowledge cutoff
- 2026-02-03
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.50
- Output token cost
- $1.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Qwen3 Coder Next FP8
No articles yet. Fetch the latest news to show it here.