Scaleway
PrismML released Ternary Bonsai 2 27B, a ternary-weight derivative of Qwen3.8 27B that compresses the 53.80 GB FP16 model to just 5.93 GB while retaining 98.2% of the parent's performance across 20 benchmarks. The 27.36B-parameter architecture keeps Qwen3.8 27B's hybrid attention backbone (roughly 75% linear, 25% full) and supports text plus image input at a 262K-token context length. Released under Apache 2.0, the quantized model targets local deployment on a 16 GB laptop or single RTX 5090 GPU, requiring PrismML's llama.cpp fork or MLX runtime. PrismML demonstrates it driving Cline coding agents and computer-use workloads locally, positioning the ternary variant as a practical quantization step two months after the original Bonsai 27B, which retained about 95% of Qwen3.8 27B.