Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Ternary Bonsai 2 27B

The model overview is temporarily unavailable.

OpenRouterprism-ml/ternary-bonsai-2-27b

Quick Info

Powered by
Provider
OpenRouter
Model key
prism-ml/ternary-bonsai-2-27b
Release date
Sep 18, 2026
Last updated
Sep 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.075
Output token cost
$0.50

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Ternary Bonsai 2 27B

OpenRouter

Coverage

Kaitchup's independent Substack review covers PrismML's release of Ternary Bonsai 2 27B as a heavily compressed version of Qwen3.8 27B at roughly 5.9 GB of language-model weights. The author notes that the first Bonsai (3.9 GB, 1-bit) was a remarkable compression result but not production-recommended: with thinking ena The piece frames Bonsai 2 27B as the subject of fresh agentic-coding evaluation, asking whether the model can finish difficult tasks, how many tokens it needs, and whether it can maintain a coding workflow without exhausting the context window. The author, Benjamin Marie, publishes weekly Kaitchup newsletters evaluatin

OpenRouter

CoverageRelease Notes

MarkTechPost reports that Ternary Bonsai 2 27B is a ternary-weight version of Qwen3.8 27B occupying 5.93 GB versus 53.80 GB in FP16, retaining 98.2% of the parent model's average across 20 benchmarks. The model accepts text and images, supports a 262K-token context, and ships under Apache 2.0 with weights runnable on a The article details the architecture: 27.36B total parameters split into a 24.35B language backbone, 2.54B embeddings and LM head, and a 0.47B vision tower, with hybrid attention at roughly 75% linear and 25% full-attention layers. Only 26.2M parameters (0.0976%) stay in higher precision, covering the recurrent state p

OpenRouter

Coverage

DataCamp's explainer covers the September 17, 2026 release of Ternary Bonsai 2 27B, a compressed version of Qwen3.8 27B at 5.9 GB that retains 98.2% of aggregate benchmark performance. The compression uses ternary weights limited to −1, 0, and +1 paired with FP16 group-wise scaling, working out to 1.72 true bits per we The article reports that the model runs locally on NVIDIA GPUs via CUDA and Apple devices via MLX, reaching 143 tokens/second on an RTX 5090, and ships under Apache 2.0 with weights on Hugging Face, with third-party providers hosting it for free. DataCamp highlights that math and coding hold at parity while knowledge a

OpenRouter

Coverage

PrismML announced Bonsai 2 27B on September 17, 2026 as its new flagship ternary-quantized model derived from Qwen3.8 27B. The release targets reasoning, coding, vision, and agentic capabilities while drastically reducing deployment size. According to the announcement, Ternary Bonsai 2 27B occupies 5.9 GB, achieving mo PrismML founder and CEO Babak Hassibi framed the model as closing the quality gap for local deployment, highlighting agentic coding, multimodal understanding, and long-horizon task execution as supported workflows. Ion Stoica, PrismML advisor and UC Berkeley professor, commented on the significance of retaining capabil

Videos about Ternary Bonsai 2 27B