Currently listed through these providers:
Model details
Qwen3 235B A22B
Qwen3 235B A22B is a Mixture-of-Experts model in the Qwen3 family that pairs 235 billion total parameters with an active parameter footprint of 22 billion, a design choice that aims to deliver strong reasoning quality while keeping per-token compute comparatively modest. Independent coverage of the closely related Instruct-2507 release describes this sparse layout and confirms that the model ships in both a standard precision and an FP8 quantized variant, with the FP8 build specifically tuned to reduce memory usage during inference without visibly sacrificing output quality. That quantization framing makes the FP8 deployment especially attractive for serving scenarios where hosting costs or accelerator memory are the limiting factor, since it lets operators fit the same expert routing behaviour into leaner hardware budgets while still tapping into the full 235B-parameter knowledge base at need.
Practical fit for this model centers on long, information-dense workloads. The underlying Qwen3-235B-A22B family is reported to support a native context of roughly 262,144 tokens, which opens up use cases like whole-document analysis, multi-file code review, lengthy conversational memory, and research-style synthesis where smaller-context models tend to truncate or summarize away important detail. Outside of pure language tasks, the same family has been positioned for multilingual coverage across more than a hundred languages as well as reasoning over visual and audio inputs alongside tool use, suggesting the FP8 build is best suited to teams that want a single backbone for mixed retrieval, reasoning, and code or agent-style workflows. It is also documented for inference on the AWS Neuron stack, which broadens the deployment options for organizations already invested in that hardware path and signals ongoing ecosystem support for running this MoE architecture at production scale.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- qwen/qwen3-235b-a22b-fp8
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.80
Limits
- Output tokens
- 20,000 tokens
- Context window
- 40,960 tokens
Latest news about Qwen3 235B A22B
No articles yet. Fetch the latest news to show it here.