Lucida Real-to-Sim

NEW

Key Features

Lucida follows a parse, generate, and place pipeline.
Reference views, masks, partial point clouds, boxes, and referring cues guide reconstruction.
Asset generation targets complete objects despite occlusion in the original capture.
GizmoAct is a vision-language agent that operates a 3D editor in a closed loop.
Executable edits adjust object poses using feedback from rendered observations.
The resulting scene contains independent editable mesh assets.
The project studies object detection, pose estimation, and scene reconstruction.
Reported evaluations include R2S, CA-1M, and Aria Digital Twin.

The pipeline gathers instance evidence from posed observations, generates occlusion-free object images and 3D assets, and then refines placement with GizmoAct. This vision-language agent receives rendered views and issues executable edits in each object local coordinate frame until the reconstruction aligns with the capture.


Lucida is useful for research in real-to-simulation conversion and object-level scene editing. Its closed-loop refinement can correct errors left by earlier perception or asset generation stages. The public project provides method descriptions, examples, and evaluations; public deployment code and service pricing are not established by the supplied page.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!