Key Features

Prediction uses one Transformer pass between VAE encoding and decoding.
The model family includes surface-normal estimation.
An albedo variant predicts surface color with lighting effects removed.
The see-through depth model targets the final depth layer behind transparent surfaces.
The primary implementation adapts Qwen-Image-Edit-2509 with four-bit LoRA training.
Semantic feature alignment and Sinkhorn-based matching improve thin structures and depth discontinuities.
A test-time adapter and scale-shift fitting incorporate sparse metric measurements.
The project links public weights, source code, and an interactive demonstration.

The main recipe adapts a pretrained image-editing diffusion Transformer into a single-step predictor. Four-bit quantization and LoRA reduce training requirements, while semantic representation alignment and a second-stage Sinkhorn-based loss improve structure and boundaries. Different supervision targets produce different dense prediction models.


Marigold V2 is useful for practitioners seeking detailed predictions without iterative diffusion sampling. The project reports training within a week on one consumer GPU and inference at 2048 by 2048 on a 32GB GPU. Code, weights, and a separate interactive demo are linked from the supplied Space.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!