Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NEAR AI Cloud logo

Model details

Qwen 3.6 35B A3B FP8

Hosted on NEAR AI Cloud as part of the Qwen family, this variant is cataloged as a chat-type model with 36B parameters and an unusually large 262,144-token context window, giving it room to hold entire repositories and lengthy technical discussions in memory during a single session. Aggregator descriptions position it as tuned for agentic coding workflows, with fluency in frontend tasks and repository-level reasoning, making it a practical fit for developer assistants that need to navigate multi-file codebases rather than just answer isolated questions. The FP8 quantization in the name suggests a precision-optimized build aimed at balancing quality with efficient inference on modern GPU fleets.

The cataloged capability set includes function calling, structured output, and extended thinking, which together support tool-augmented agents that need to reason carefully before acting, as well as streaming for responsive interactive use. Third-party benchmark figures reported by Inferbase include GPQA at 95.3, AA-LCR at 84.1, HLE at 63.7, and Tau2-Bench at 64.4, alongside a Tau2 Banking score of 41.9 and a Terminal-Bench 2.1 score of 44.9, indicating strong general reasoning performance alongside more mixed results on specialized agentic and banking benchmarks. With its open-weights posture and very wide context window, the model is well suited for teams building long-horizon coding assistants, document-grounded chatbots, and pipeline-friendly agent prototypes that want open deployment flexibility.

NEAR AI CloudQwen/Qwen3.6-35B-A3B-FP8qwen

Quick Info

Powered by
Provider
NEAR AI Cloud
Model key
Qwen/Qwen3.6-35B-A3B-FP8
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.17
Output token cost
$1.10

Limits

Output tokens
8,192 tokens
Context window
262,144 tokens

Latest news about Qwen 3.6 35B A3B FP8

NEAR AI Cloud

CoverageRelease Notes

Venice.ai's published changelog lists "Qwen 3.6 35B A3B" as a Text model added in its "New Models" section, classified under the "Private" privacy mode, made available to all users, with a stated rollout date of July 20, 2026. The entry appears within the broader May 26, 2026 – July 27, 2026 change-log window that also Because the Venice.ai entry names the model without the "FP8" suffix present in the NEAR AI Cloud modelKey, the FP8 designation cannot be verified from this page alone, and the changelog documents Venice.ai's offering rather than NEAR AI Cloud's hosted endpoint at cloud-api.near.ai/v1. The page is a third-party aggrega

NEAR AI Cloud

Coverage

The official Qwen model card on ModelScope documents the Qwen3.6-35B-A3B-FP8 release, the first open-weight variant of the Qwen3.6 series, shipped as fine-grained FP8-quantized weights (block size 128) with performance metrics described as nearly identical to the original model. The card states the artifact is compatib Architecture details confirm the exact subject: a 35B-parameter causal language model with vision encoder, 3B activated parameters per token, 40 layers arranged as 10 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)), 256 experts with 8 routed plus 1 shared, hidden dimension 2048, and 262,144-token native co

Videos about Qwen 3.6 35B A3B FP8

More models around Qwen 3.6 35B A3B FP8