Key Features

Runs a 27B-class multimodal model in low-bit form on local devices.
Offers a ternary 27B variant at about 5.9 GB for laptop-class quality.
Offers a 1-bit 27B variant at about 3.9 GB for phone-class memory budgets.
Uses low-bit representation across embeddings, attention, MLPs, and the LM head.
Includes a compact 4-bit vision tower for image, document, screenshot, and camera workflows.
Supports a 262K-token context window.
Supports speculative decoding for draft-and-verify acceleration.
Ships model weights under the Apache 2.0 license with local and API access paths.

The ternary variant uses 1.71 effective bits per weight and targets laptop-class quality at about 5.9 GB, while the 1-bit variant uses 1.125 effective bits per weight and fits around 3.9 GB. Both variants keep the low-bit representation across the language network, support a compact 4-bit vision tower, offer a 262K-token context, and support speculative decoding.


Bonsai 27B is useful for private local assistants, offline agentic workflows, hybrid cloud-local systems, on-device vision use cases, and cost-sensitive agent loops. PrismML reports strong retention of full-precision baseline performance across math, coding, tool calling, instruction following, knowledge, and vision benchmarks, with Apache 2.0 weights available on release.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner
Zero to AI Engineer Program

Zero to AI Engineer

Skip the degree. Learn real-world AI skills used by AI researchers and engineers. Get certified in 8 weeks or less. No experience required.

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!