A frozen behavior policy generates trajectories and low-noise queries. Normalized reward ascent and descent form bounded positive and negative clean-output targets. The trainable policy fits detached targets for a limited update budget, after which an exponential moving average refreshes the behavior policy and the process repeats.
The public implementation supports SD3.5-M and the few-step Z-Image-Turbo regime, with reward-specific and joint training experiments. DiffusionOPSD is useful for researchers studying efficient post-training and alignment of image generators. Base-model access, reward-model dependencies, and the licenses of those components remain separate from the training code.

