A new version of a large language model, called Ternary Bonsai 2 27B, has been introduced by PrismML. This model is a compressed version of the Qwen3.8 27B, a large model developed by Alibaba Cloud. The compressed version uses only 5.9 GB of storage, which is nine times smaller than the original model, yet it retains 98.2% of the original model's performance in various tasks such as reasoning, coding, and image recognition. This makes it more practical for use on local devices like smartphones, laptops, and other edge devices instead of relying solely on cloud-based systems. PrismML is a startup focused on creating highly compressed language models that can run efficiently on local devices. The Ternary Bonsai 2 27B model uses a special compression method that represents the model's weights using only three values: -1, 0, and +1. This allows for a more compact model without losing too much performance. The model also supports both text and image inputs and has a large context window of 262,000 tokens, meaning it can process long sequences of text or images effectively. It is available under the Apache 2.0 license, which allows for free use and modification. Compared to the first version of Bonsai 27B, the new model is based on a more powerful base model, Qwen3.8 27B, and retains a higher percentage of the full-precision model’s capabilities. The improvements focus on enhancing performance in specific areas such as reasoning, programming, vision, and long-term tasks. These areas are especially important for applications that require accurate and consistent performance over time, such as coding assistants or systems that interact with the environment through agents. The Ternary Bonsai 2 27B model is designed to be energy-efficient and fast. It can generate up to 143 tokens per second on an NVIDIA GeForce RTX 5090 and 46.8 tokens per second on an Apple M5 Max. It also consumes less energy, making it more efficient for tasks that require frequent processing. This efficiency and performance make it suitable for a range of applications, from coding and debugging to analyzing private documents and images locally, reducing the need to rely on cloud resources.