Currently listed through these providers:
Model details
Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is positioned as an open-weight sparse mixture-of-experts model in the Qwen family, activating 95 billion of its 2.4 trillion total parameters per token. Alibaba released the open weights, framing it as the open-weights counterpart to Qwen3.8-Max and aiming to bring near-frontier capability into the open ecosystem. The architecture is designed for heavy lifting tasks: coding, research, complex reasoning, and agentic workflows where a single response can chain tool use, multi-step planning, and long-context synthesis.
In practical deployment, the model has already been documented for serving on NVIDIA GB300 NVL72 hardware with a configurable reasoning mode, signaling that operators can dial reasoning effort up or down depending on the workload. DeepInfra lists a hosted endpoint for text generation under the Qwen namespace, giving developers an inference option without standing up their own GPU fleet. The combination of open weights, very large total capacity with moderate per-token activation, and configurable reasoning points to a model aimed at teams that want frontier-style behavior but retain the ability to self-host, fine-tune, and control cost on large-context jobs.
Quick Info
Powered by- Provider
- Charm Hyper
- Model key
- qwen3.8-2.4t-a95b
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.00
- Output token cost
- $6.00
Limits
- Output tokens
- 128,000 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.8 2.4T A95B
Videos about Qwen3.8 2.4T A95B
More models around Qwen3.8 2.4T A95B
This exact model name is also listed by 12 other providers.