Alibaba's Qwen3.8 Max is a sparse mixture-of-experts model with roughly 2.4 trillion total parameters and about 95 billion active per forward pass, an architecture that keeps inference tractable while supporting a very large overall capacity. It was first exposed through Alibaba's QwenCloud API, then published as the open-weight Qwen3.8-2.4T-A95B checkpoint on Hugging Face and ModelScope under an Apache 2.0 license, giving well-resourced teams the option to self-host rather than rent access. A 27-billion-parameter dense sibling shipped open weights a day later for teams that cannot provision hardware for the full MoE, while the hosted Qwen3.8-Max endpoint remains the route for reaching the larger configuration directly.
Qwen3.8 Max is positioned for software engineering and long-horizon agent work rather than general chat. On Terminal-Bench 2.1 it lands within roughly two points of top closed frontier systems, and on the updated Qwen3.8-Max-0902 build it closes additional ground on agent-focused suites such as DeepSWE and MLS-Bench-Lite. Independent benchmark tracking gives the model a top-five placement for multimodal and grounded tasks and a top-ten placement for agentic tool use, with notably strong performance on screenshots, charts, and document comprehension. That balance of coding, reasoning, and multimodal grounding makes it a natural fit for coding agents and retrieval-augmented workflows where a self-hostable alternative to closed frontier APIs is desired.