The pipeline gathers instance evidence from posed observations, generates occlusion-free object images and 3D assets, and then refines placement with GizmoAct. This vision-language agent receives rendered views and issues executable edits in each object local coordinate frame until the reconstruction aligns with the capture.
Lucida is useful for research in real-to-simulation conversion and object-level scene editing. Its closed-loop refinement can correct errors left by earlier perception or asset generation stages. The public project provides method descriptions, examples, and evaluations; public deployment code and service pricing are not established by the supplied page.

