Key Features

Fully open Mixture-of-Experts language model from AMD.
Uses 16B total parameters with about 2.8B active parameters per token.
Trained from scratch on AMD Instinct MI300X and MI325X GPUs.
Releases checkpoints across pretraining, mid-training, SFT, DPO, and RL stages.
Provides Hugging Face model links for multiple training checkpoints.
Includes a public GitHub repository for code and technical assets.
Designed for efficient MoE inference and open LLM experimentation.
Documents model architecture, training pipeline, and benchmark positioning.

The release covers the full training pipeline, including pretraining, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning. AMD provides multiple checkpoints so researchers can inspect and reuse different stages rather than only a final instruction model.


Instella-MoE is useful for open LLM research, efficient inference, training-pipeline study, and AMD GPU optimization. Its combination of public code, public checkpoints, and documented training stages makes it valuable for teams that need reproducible MoE model development.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!