The family uses a unified diffusion framework and builds its visual prior through image-only pretraining before paired language supervision and joint generation-editing training. The Base model uses a longer sampling schedule, while the Turbo model applies Twin-DMD distillation for generation in a few steps.
LLaDA-Image is useful for developers building local generation and editing workflows with shared infrastructure. Public releases include Base and Turbo checkpoints, BF16 and FP8 variants, and Diffusers-based inference code. The repository identifies training code as forthcoming, so the available release should not be described as a complete training pipeline.

