Key Features

The method post-trains diffusion image-generation models.
Reward ascent and descent construct bounded positive and negative clean-output targets.
The behavior policy supplies queries at low-noise points along generated trajectories.
The trainable policy fits detached targets under a finite update budget.
An exponential moving average updates it between training iterations.
The release includes SD3.5-M and Z-Image-Turbo configurations.
Experiments include reward-specific and joint training settings.
Explicit targets allow separate analysis of target quality and realized policy improvement.

A frozen behavior policy generates trajectories and low-noise queries. Normalized reward ascent and descent form bounded positive and negative clean-output targets. The trainable policy fits detached targets for a limited update budget, after which an exponential moving average refreshes the behavior policy and the process repeats.


The public implementation supports SD3.5-M and the few-step Z-Image-Turbo regime, with reward-specific and joint training experiments. DiffusionOPSD is useful for researchers studying efficient post-training and alignment of image generators. Base-model access, reward-model dependencies, and the licenses of those components remain separate from the training code.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!