Key Features

Expert weights are memory-mapped and streamed from SSD only when needed.
A trained prediction head schedules upcoming expert loads ahead of the forward pass.
Distilled adapters help recover quality lost when the base model is quantized to four bits.
The current MLX backend supports macOS on Apple Silicon.
The release includes Edge0-35B-A3B and Edge0-8B-A1B preview checkpoints.
Each released model directory includes its matching LoRA and prerouter weights.
Backend isolation lets implementations share the same core inference abstractions.
CUDA is a planned backend rather than a currently supported platform.

The framework combines on-demand expert loading, a trained prerouter that predicts upcoming expert selections, and Recover-LoRA adapters distilled from a full-precision teacher. Released model tiers bundle the checkpoint, adapters, and prerouter together. Backend abstractions separate portable inference logic from the MLX implementation.


Edge0 suits developers experimenting with local MoE serving on compatible Macs. It ships 35B-A3B and 8B-A1B preview tiers and documents memory, disk, and software requirements. CUDA is on the roadmap; the current supported backend requires macOS and Apple Silicon.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!