PrismML has released Ternary Bonsai 2 27B, a multimodal model based on Qwen3.8 27B and designed for local deployment. It uses ternary weights of -1, 0, and +1 with FP16 group-wise scaling, achieving 1.76 effective bits per weight and a total footprint of 5.9GB. The model supports a 262K-token context window, text-and-image input, and the Apache 2.0 license. PrismML reports an overall benchmark score of 83.9, or 98.2% of the full-precision Qwen3.8 27B score of 85.4; the evaluation covers reasoning, mathematics, coding, instruction following, vision, and agentic tool use. The reported retention varies by capability: math reaches 96.57 versus 97.06 for the full-precision model, while coding scores 81.58 versus 82.17 and vision scores 78.59 versus 81.64. Compared with the previous Bonsai 27B generation, the company says the new model improves its base model, reasoning, coding, vision, and long-horizon agentic performance, raising retention from 95% to more than 98%. PrismML also reports throughput of up to 143 tokens per second on an NVIDIA GeForce RTX 5090 and 46.8 tokens per second on an Apple M5 Max. On an RTX 4090, it consumes 0.714 mWh per token, which the company says is 40% more energy-efficient than an 8B model running at full precision. The model runs on NVIDIA GPUs through CUDA and on Apple devices through MLX using custom low-bit kernels, with weights available immediately under Apache 2.0. The release is positioned for coding agents, computer use, private document analysis, multimodal debugging, and hybrid workflows that keep selected tasks local.
