The main recipe adapts a pretrained image-editing diffusion Transformer into a single-step predictor. Four-bit quantization and LoRA reduce training requirements, while semantic representation alignment and a second-stage Sinkhorn-based loss improve structure and boundaries. Different supervision targets produce different dense prediction models.
Marigold V2 is useful for practitioners seeking detailed predictions without iterative diffusion sampling. The project reports training within a week on one consumer GPU and inference at 2048 by 2048 on a 32GB GPU. Code, weights, and a separate interactive demo are linked from the supplied Space.

