Currently listed through these providers:
Model details
Qwen3.6 27B FP8
As a member of Alibaba's Qwen3.6 family, the 27B parameter model is built around practical multimodal chat and agent workflows, with native text, image, video, and audio inputs that produce text responses. The FP8 quantization preserves the architecture's full reasoning and instruction-following behavior while lowering memory and compute demands, making the weights easier to host on a single high-end accelerator such as an H100. Distributed under Apache 2.0, the checkpoint is intended for teams that want a mid-sized open model they can self-deploy, fine-tune, or wrap behind their own inference stack without proprietary licensing constraints.
The design sweet spot is long-context reasoning paired with tool use and structured outputs, letting applications pass large document bundles, transcripts, or multi-turn histories while still invoking external APIs and producing schema-conformant responses. FP8 inference keeps latency and per-token energy reasonable for sustained workloads, and the open weights let developers inspect the model, adapt prompts, or distill downstream variants. Versus the family's heavier 35B-A3B MoE sibling, this 27B dense build favors simpler serving, predictable cost, and easier integration for teams that want strong general reasoning without operating a sparse-expert deployment.
Quick Info
Powered by- Provider
- InferX
- Model key
- Qwen3.6-27B-FP8
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens