Currently listed through these providers:
Model details
Qwen 3.5 9B (MLX 4-bit)
As a 4-bit MLX quantization of the Qwen 3.5 family, this release is positioned for efficient on-device and developer-machine inference rather than for raw frontier scale. The technical evidence lists it at roughly 1.9B parameters under the qwen3.5 architecture, which is noticeably smaller than the 9B suggested by the model name and worth keeping in mind when weighing capability against efficiency. It is distributed under the Apache-2.0 license and is openly available through a community Hugging Face repository, reinforcing its appeal for teams that want a permissive, portable base model they can fine-tune or inspect.
In practical terms, the endpoint accepts both text and image inputs while returning text, and it supports structured interactions through tool calling, attachment handling, and adjustable sampling. The serving profile is generous for its size, with a 33K-token context window and up to 8K tokens of output per response, making it well suited to moderately long document analysis, code review, and image-grounded Q&A where larger frontier models would be overkill. Because pricing is set at zero on both input and output, the deployment is essentially a no-cost sandbox for prototyping multimodal assistants or evaluating Qwen 3.5 behavior before committing to heavier production workloads.
Quick Info
Powered by- Provider
- Atomic Chat
- Model key
- Qwen3_5-9B-MLX-4bit
- Release date
- Mar 5, 2026
- Last updated
- Apr 4, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 32,768 tokens
Latest news about Qwen 3.5 9B (MLX 4-bit)
No articles yet. Fetch the latest news to show it here.
