A compressed version of the Qwen3.8-27B model from Alibaba, named Bonsai 2 27B, was tested on four different machines with varying hardware and software setups. This version, reduced to 5.95 GB by an American company called Prism ML, became available on September 17, 2026, under the Apache 2.0 license. The model weights can be downloaded from Hugging Face, but they require a special runtime from Prism ML, which isn't compatible with common tools like LM Studio or Ollama. The four tested machines included an Intel mini PC with a Core Ultra processor and Arc graphics, a Mac mini M6 with 24 GB of unified memory, a gaming PC with an RTX 4070 Ti SUPER graphics card, and an AMD mini PC based on the Ryzen AI Max+ 395. All used the same version of the Prism ML runtime and the same 5.95 GB Bonsai file. A Qwen3.8 file in Q4 format was also used for comparison purposes. The performance of the models was measured using the llama-bench tool, with a 512-token prompt followed by a 128-token generation, repeated five times. The RTX 4070 Ti SUPER gaming PC performed best with Bonsai, achieving nearly 69 tokens per second. The Mac mini M6 managed 18.1 tokens per second, while the Intel mini PC was much slower, at 1.75 tokens per second. The performance of Bonsai versus Q4 varied depending on the machine and the backend used. Bonsai outperformed Q4 on CUDA (Nvidia’s backend) and Metal (Apple’s backend), but Q4 performed better under Vulkan (Intel’s backend). The AMD mini PC showed a significant improvement when using ROCm instead of Vulkan, with Bonsai’s throughput increasing more than five times. The RTX 4070 Ti SUPER’s performance was also affected by the amount of VRAM available. When the Q4 model couldn’t fit entirely into the GPU’s memory, parts had to be stored in the system memory, slowing down processing. Moving some of the model’s layers to the CPU significantly improved the throughput. The tests showed that AI model speed depends not only on hardware but also on the available kernels for a specific format and the software backend used. While the quality of Bonsai wasn’t evaluated in these tests, Prism ML claims it retains about 98.2% of the original performance. However, some real-world tests showed the model might miss certain details or have issues with specific tasks. The best machine for running local AI depends on the use case. The RTX 4070 Ti SUPER was the most responsive with Bonsai but could be limited by its VRAM. The Mac mini M6 offered a quiet, balanced option, while the AMD mini PC, with its large unified memory, showed the highest performance under ROCm. The Intel mini PC was the slowest of the four for this use. The tests also highlighted the importance of software in AI performance, as changing the backend could significantly affect throughput. Future improvements in Vulkan could potentially change the current performance landscape.