Currently listed through these providers:
Model details
Qwen3 Coder Next FP8
Qwen3 Coder Next FP8 is an open-weight language model purpose-built for coding agents and local development, with a focus on keeping response times fast while still handling the kind of multi-step, tool-driven work that real software tasks demand. Rather than relying on a large number of activated parameters for every pass, the design uses a mixture-of-experts setup with 80B total parameters but only 3B activated per token. This efficiency story is reinforced in third-party write-ups, which describe performance comparable to much denser models with ten to twenty times more active parameters, giving developers a cost-effective foundation for agent deployments that need to scale across many concurrent tasks.
Beyond raw efficiency, the model leans into practical agentic capability through its long context window, which integrates well with popular CLI and IDE environments such as Claude Code, Qwen Code, Qoder, Kilo, Trae, and Cline. It is positioned as a non-reasoning, ultra-quick coding model that excels at long-horizon reasoning, complex tool usage, and recovering from execution failures during dynamic tasks, making it a strong fit for scaffolding around real codebases rather than just single-shot completions. The FP8 checkpoint itself applies fine-grained FP8 quantization with a block size of 128 to make inference lighter, while the published benchmarks come from the original bfloat16 model so accuracy claims remain tied to the unquantized behavior, and dynamic quantized variants from the community further broaden the hardware options for local and self-hosted setups.
Quick Info
Powered by- Provider
- InferX
- Model key
- qwen3-coder-next-fp8
- Release date
- Feb 3, 2026
- Last updated
- Feb 3, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 65,536 tokens
- Context window
- 256,144 tokens