Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Together AI logo

Model details

Qwen3 235B A22B Instruct 2507 FP8

The Qwen3 235B A22B Instruct 2507 FP8 represents Alibaba Cloud's effort to deliver massive scale within practical deployment constraints. This model uses a mixture-of-experts architecture with 128 specialized expert networks, though only 8 activate during any given inference step, keeping active computation at 22 billion parameters despite the full 235-billion-parameter scale. The architecture spans 94 layers with grouped query attention, optimized through 8-bit floating-point quantization to enable faster processing without sacrificing output quality. Native support for 262K token contexts makes it well-suited for analyzing lengthy documents, conducting extended multi-turn conversations, and handling complex agentic workflows that require maintaining context across thousands of tokens.

This instruction-tuned variant emerged from the Qwen3 family with targeted post-training enhancements focused on real-world utility. The update delivers measurable improvements across instruction following, logical reasoning, mathematical problem-solving, scientific comprehension, and code generation, while also excelling at tool-calling tasks. User preference alignment received particular attention during refinement, yielding more helpful and higher-quality responses in open-ended and subjective scenarios. The model's coverage of long-tail knowledge across multiple languages supports multilingual applications, and its Apache 2.0 licensing opens both research and commercial deployment paths. Performance benchmarks position it competitively—achieving notably strong results on reasoning and mathematics evaluations—making it a strong choice for teams needing powerful open-weight inference without proprietary model restrictions.

Together AIQwen/Qwen3-235B-A22B-Instruct-2507-tputqwendeprecated

Quick Info

Powered by
Provider
Together AI
Model key
Qwen/Qwen3-235B-A22B-Instruct-2507-tput
Release date
Jul 25, 2025
Last updated
Jul 25, 2025
Knowledge cutoff
2025-07
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Qwen3 235B A22B Instruct 2507 FP8

No articles yet. Fetch the latest news to show it here.

Videos about Qwen3 235B A22B Instruct 2507 FP8

More models around Qwen3 235B A22B Instruct 2507 FP8