The framework combines on-demand expert loading, a trained prerouter that predicts upcoming expert selections, and Recover-LoRA adapters distilled from a full-precision teacher. Released model tiers bundle the checkpoint, adapters, and prerouter together. Backend abstractions separate portable inference logic from the MLX implementation.
Edge0 suits developers experimenting with local MoE serving on compatible Macs. It ships 35B-A3B and 8B-A1B preview tiers and documents memory, disk, and software requirements. CUDA is on the roadmap; the current supported backend requires macOS and Apple Silicon.

