Sulat.com
AI models
InferX logo

Model details

Qwen3.6 27B FP8

As a member of Alibaba's Qwen3.6 family, the 27B parameter model is built around practical multimodal chat and agent workflows, with native text, image, video, and audio inputs that produce text responses. The FP8 quantization preserves the architecture's full reasoning and instruction-following behavior while lowering memory and compute demands, making the weights easier to host on a single high-end accelerator such as an H100. Distributed under Apache 2.0, the checkpoint is intended for teams that want a mid-sized open model they can self-deploy, fine-tune, or wrap behind their own inference stack without proprietary licensing constraints.

The design sweet spot is long-context reasoning paired with tool use and structured outputs, letting applications pass large document bundles, transcripts, or multi-turn histories while still invoking external APIs and producing schema-conformant responses. FP8 inference keeps latency and per-token energy reasonable for sustained workloads, and the open weights let developers inspect the model, adapt prompts, or distill downstream variants. Versus the family's heavier 35B-A3B MoE sibling, this 27B dense build favors simpler serving, predictable cost, and easier integration for teams that want strong general reasoning without operating a sparse-expert deployment.

InferXQwen3.6-27B-FP8qwen

Quick Info

Powered by
Provider
InferX
Model key
Qwen3.6-27B-FP8
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen3.6 27B FP8

Videos about Qwen3.6 27B FP8

More models around Qwen3.6 27B FP8