Key Features

14 MB model binary.
Runs in about 28 MB of RAM.
Tool calling and structured extraction.
Device-use and mobile-agent workflows.
CQ2-bit Cactus Quants compression.
WebAssembly browser execution.
Parallel and multi-step tool calls.
High on-device token throughput.

It is built on the Simple Attention Network, compressed to CQ2-bit with Cactus Quants, and packaged in Cactus's own engine. The browser demo runs the model in WebAssembly and shows tool routing, parallel calls, multi-step actions, extraction, classification, and refusal behavior.


Needle 2 is useful for offline assistants, smart-home control, phones, wearables, robots, and embedded interfaces where memory, latency, and connectivity are constrained. Cactus reports up to 500 tokens per second on a Raspberry Pi 5 and high throughput on supported XR devices and phones.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!