Bonsai 2 27B is a highly compressed 27B-class multimodal model derived from Qwen3.8 27B that uses ternary weights {−1, 0, +1} with FP16 group-wise scaling to achieve 1.76 effective bits per weight and a 5.9 GB total footprint. It supports a 262k-token context window and multimodal text-and-image input, and is released under the Apache 2.0 license. Despite being more than 9× smaller than its full-precision counterpart, it retains 98.2% of aggregate benchmark performance (overall score 83.9 versus Qwen3.8 85.4) and preserves capability in sensitive areas like coding, long-horizon agentic workflows, vision, reasoning, and instruction following.
The release emphasizes deployment gains: up to 143 tokens/sec on an NVIDIA GeForce RTX 5090, 46.8 tokens/sec on Apple M5 Max, and 0.714 mWh/token on an RTX 4090 (about 40% more energy-efficient than an 8B full-precision model). Custom low-bit kernels enable execution on CUDA GPUs and Apple devices. By closing the retention gap from earlier Bonsai releases (95% to >98%), the model makes near-lossless low-bit deployment practical for local assistants, private document analysis, coding agents, and hybrid local/cloud architectures. Full technical details, weights, demos, and partnership options are provided, positioning this model as a path to greater intelligence density and lower operational cost across devices and datacenters.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.