Currently listed through these providers:
Model details
Qwen3.5 122B-A10B
Qwen3.5 122B-A10B sits within the Qwen model family and follows the lineage of large, long-context Qwen releases that have been explored by open-source users for both hosted and local deployment. Community interest has focused on running the model on compact single-node hardware, including a RedHatAI-distributed NVFP4 quantization that NVIDIA Developer Forums users identified as a workable option for a single DGX Spark / GB10 setup. A follow-up forum thread further reported an on-device throughput figure of up to 51 tokens per second in a v2.1 release that bundled patches, a quick-start guide, and benchmark notes, reflecting active community effort to optimize the model for constrained environments.
Practically, the model is positioned for text-based workloads where long context handling is useful, and the presence of open-weight community quantizations like the NVFP4 build makes it accessible to teams that want to self-host on enthusiast-class GPU hardware. The forum activity around quick-start guides and patches suggests the model rewards hands-on tuning and is being adopted by technically engaged users who need to balance capability with the limits of a single workstation. Teams evaluating Qwen3.5 122B-A10B should weigh the demonstrated community interest in local inference against the level of engineering effort required to stabilize the model on their target hardware.
Quick Info
Powered by- Provider
- Neon
- Model key
- qwen35-122b-a10b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.22
- Output token cost
- $2.20
Limits
- Output tokens
- 25,000 tokens
- Context window
- 262,144 tokens