Key Features

Fire3D accepts an unsegmented RGB image or casual RGB video.
The model predicts object decomposition without supplied masks, boxes, or prompts.
The output includes object-level six-degree-of-freedom poses and oriented boxes.
Amodal reconstruction fills unseen geometry rather than only meshing visible surfaces.
The generated assets include material and PBR appearance information.
Foreground entities remain separate, transformable mesh assets.
The main reconstruction system is feed-forward without test-time scene optimization.
The repository lists training and inference code, model checkpoints, and processed examples.

The feed-forward pipeline lifts visual features into 3D, predicts instance structure, and uses canonicalized object evidence to condition shape and material generation. A highly compressed VAE supports batched processing of objects, after which meshes are placed back into the estimated metric scene layout.


Fire3D is designed for scene reconstruction and simulation-data workflows that need actual object meshes. The project reports reconstruction in under a minute and releases inference code, training code, models, and example inputs. Its reference protocols use substantial GPU memory, so local deployment requires checking the documented environment.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!