Hosted on NEAR AI Cloud as part of the Qwen family, this variant is cataloged as a chat-type model with 36B parameters and an unusually large 262,144-token context window, giving it room to hold entire repositories and lengthy technical discussions in memory during a single session. Aggregator descriptions position it as tuned for agentic coding workflows, with fluency in frontend tasks and repository-level reasoning, making it a practical fit for developer assistants that need to navigate multi-file codebases rather than just answer isolated questions. The FP8 quantization in the name suggests a precision-optimized build aimed at balancing quality with efficient inference on modern GPU fleets.
The cataloged capability set includes function calling, structured output, and extended thinking, which together support tool-augmented agents that need to reason carefully before acting, as well as streaming for responsive interactive use. Third-party benchmark figures reported by Inferbase include GPQA at 95.3, AA-LCR at 84.1, HLE at 63.7, and Tau2-Bench at 64.4, alongside a Tau2 Banking score of 41.9 and a Terminal-Bench 2.1 score of 44.9, indicating strong general reasoning performance alongside more mixed results on specialized agentic and banking benchmarks. With its open-weights posture and very wide context window, the model is well suited for teams building long-horizon coding assistants, document-grounded chatbots, and pipeline-friendly agent prototypes that want open deployment flexibility.