The ternary variant uses 1.71 effective bits per weight and targets laptop-class quality at about 5.9 GB, while the 1-bit variant uses 1.125 effective bits per weight and fits around 3.9 GB. Both variants keep the low-bit representation across the language network, support a compact 4-bit vision tower, offer a 262K-token context, and support speculative decoding.
Bonsai 27B is useful for private local assistants, offline agentic workflows, hybrid cloud-local systems, on-device vision use cases, and cost-sensitive agent loops. PrismML reports strong retention of full-precision baseline performance across math, coding, tool calling, instruction following, knowledge, and vision benchmarks, with Apache 2.0 weights available on release.


