Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Infomaniak logo

Model details

Qwen3.5 122B-A10B FP8

Qwen3.5-122B-A10B-FP8 is a vision-language Mixture-of-Experts model with roughly 10B of its 122B parameters activated per token across 48 layers, quantized to FP8 via fine-grained block-wise methods to trim memory while preserving output fidelity. Its hybrid Gated Delta Network paired with Gated Attention blends linear and standard attention, and sparse MoE feed-forward layers keep inference efficient for long inputs. Training uses early-fusion multimodal methods so text, image, and video understanding share a unified representation, and the checkpoint runs in a thinking mode by default while supporting around 201 languages.

The native 256K-token context window, extendable toward one million tokens through YaRN, makes the model well suited to extended document analysis, code repositories, and agentic workflows that need sustained reasoning over very long inputs. Independent benchmark reporting places it as competitive with or ahead of leading closed peers on reasoning, coding, vision, and agentic evaluations, signaling a meaningful step forward for open-weight multimodal systems. Practically, it fits teams that want frontier-class reasoning and multimodal grounding in a deployable, memory-efficient open package, especially when paired with hardware configurations tuned for high-throughput FP8 inference.

InfomaniakQwen/Qwen3.5-122B-A10B-FP8qwen

Quick Info

Powered by
Provider
Infomaniak
Model key
Qwen/Qwen3.5-122B-A10B-FP8
Release date
Feb 23, 2026
Last updated
Aug 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$3.97

Limits

Input tokens
200,000 tokens
Output tokens
65,536 tokens
Context window
200,000 tokens

Latest news about Qwen3.5 122B-A10B FP8

Infomaniak

Coverage

vLLM Recipes publishes first-party deployment documentation for Qwen/Qwen3.5-122B-A10B, covering BF16 (293 GB), FP8 (147 GB), NVFP4 (74 GB), and INT4 (74 GB) checkpoints with a Red Hat FP8 Dynamic option. The model is described as a mid-tier Qwen3.5 family member using gated delta networks with a MoE architecture of 12 This page is the most directly relevant official technical reference for Qwen3.5-122B-A10B FP8, with explicit variant coverage and concrete deployment recipes. Model focus is satisfied because the documentation is about deploying the model itself rather than any gateway, reseller, or serving-provider pricing announceme

Infomaniak

CoverageRelease Notes

NVIDIA NIM for LLM and VLM release 2.0.13 ships optimized profiles for 13 NIMs, explicitly including qwen3.5-122b-a10b alongside other Qwen3.5 variants (qwen3.5-397b-a17b), Nemotron 3 models, Kimi K2.6, GPT-OSS 120B, Llama 3.x variants, and Gemma 4. The release delivers throughput optimizations of up to 2.98x depending This is first-party technical news from NVIDIA's NIM documentation that names the exact Qwen3.5-122B-A10B variant as receiving an officially packaged, throughput-optimized NIM profile in release 2.0.13. The model focus is satisfied because the news centers on the model itself and its optimized serving profile, not on I

Infomaniak

CoverageBenchmark

Millstone AI publishes an inference benchmark for Qwen3.5-122B-A10B-FP8 describing it as a 122B-parameter Mixture-of-Experts vision-language model with 256 experts and 10B activated parameters, quantized to FP8 via fine-grained block-wise quantization. The architecture combines a hybrid Gated Delta Network with Gated A On tested hardware (2x RTX Pro 6000 Blackwell, 192GB VRAM), the model reaches 237 tok/s peak throughput across 1K–256K context lengths at concurrency 1–5. Per-use-case capacity numbers include 4 concurrent code-completion sessions at 1K context, 102 short-form chat sessions at 8K, 7 general chatbot users at 32K, 4 docu

Videos about Qwen3.5 122B-A10B FP8

More models around Qwen3.5 122B-A10B FP8