A startup called PrismML, founded by researchers from the California Institute of Technology (Caltech) and led by Babak Hassibi, a professor and expert in compression technologies, has developed a method to shrink large language models (LLMs) significantly without sacrificing performance. On Thursday, PrismML released a compressed version of Qwen3.8 27B, a popular open-source language model from Alibaba, called Bonsai 2 27B. This compressed model is 9 to 10 times smaller in memory than the original, reducing its size to 5.9 gigabytes (GB). This makes it small enough to run on a regular personal computer and possibly even a high-end smartphone, which could bring powerful AI capabilities to more accessible devices.
The compression technique used by PrismML relies on "ternary" weights—values that are simplified from the standard 16 bits per weight used in traditional models to just three possible values: +1, -1, or 0. This simplification drastically reduces the storage space required for the model. The Bonsai 2 model performs nearly as well as the original Qwen3.8 27B, achieving 98% of its benchmark scores. This is an improvement from the previous Bonsai model, which reached 95% performance and was downloaded over 11 million times. Even smaller versions of the model have been downloaded an additional 2.6 million times, showing strong interest in the technology.
Hassibi noted that achieving 100% performance parity with the original models might not be possible, but a 2% drop in performance is likely negligible in most real-world applications. He emphasized that the environment and software surrounding a model can also greatly influence its accuracy. Looking ahead, PrismML aims to apply this compression method to even larger models, potentially those with hundreds of billions of parameters. Hassibi believes that larger models might be easier to compress without significant loss of intelligence.
Ion Stoica, an advisor to PrismML and co-founder of Databricks, praised the technology, saying it could allow advanced AI models to run directly on user devices, providing powerful intelligence locally rather than relying on cloud services. This would offer users more privacy and faster response times without the need to send data to remote servers. PrismML has received backing from investors including Khosla Ventures, Cerberus Capital, and Caltech. While other companies like Multiverse Computing are also exploring LLM compression, Hassibi claims that PrismML's method is unique in maintaining virtually no performance loss compared to the original models.
Startup Compresses Large Language Models to Run on PCs and Smartphones
AI-rewritten from original reportingHow it works
llm-compressionprismmlbonsaiai-modelscaltechternary-weights



