PrismML's Ternary Bonsai 2 27B cuts AI costs without sacrificing quality
PrismML released Ternary Bonsai 2 27B on September 17, 2026, a model based on Qwen3.8 that achieves 98.2% of its full-precision counterpart's benchmark performance in a 5.9GB footprint—9 times smaller. This compression uses ternary weights (-1, 0, +1) with FP16 group-wise scaling, enabling it to run on NVIDIA GPUs via CUDA and Apple devices via MLX. The model supports 262,000-token context windows and multimodal text-image inputs, while consuming 0.714 mWh per token on an RTX 4090—40% more energy-efficient than an 8B full-precision model. Its Apache 2.0 license ensures open access.
The ternary weights and group-wise scaling mechanism drastically reduce memory bandwidth without significant performance loss. This allows smaller devices and lower-cost infrastructure to run complex AI tasks. For knowledge access, it means more people can use advanced language models without expensive hardware or high data costs.
This advancement moves knowledge affordability by enabling cheaper, cross-platform AI services. Lower memory demands and energy use reduce barriers for low-resource users and organizations. The model’s multimodal support also expands access to diverse knowledge formats.
Real-world adoption will determine its impact. Current benchmarks are specific to listed tasks, and the model’s future performance may shift with new use cases. The release date being in the future per source also requires verification. This model shows how technical compression can expand access without quality tradeoffs—critical for global knowledge equity.
Source: Hacker News
MANY MINDED