thumbnail

InLayDiffusion: Indoor Layout Estimation from a Single Panorama via Structural Point Cloud Lifting and Guided Diffusion

Giovanni Pintore, Uzair Shah, Marco Agus, and Enrico Gobbetti

2026

Abstract

We introduce a novel end-to-end deep-learning method that combines structural geometric lifting with prior-guided diffusion to recover a 2D floorplan and a 3D room layout from a 360-degree image. Structural lifting predicts gravity-aligned, structure-aware depth and projects the resulting colored 3D point cloud onto the floor plane through a differentiable operation, directly linking panoramic image features to geometric reasoning and improving robustness to clutter while encoding cues such as room height and coarse footprint geometry. Footprint reconstruction is then formulated as polygon denoising and completion within a diffusion framework. The layout is represented as a fixed-vertex polygon jointly encoded with floor-projected RGB and density features, enabling refinement of noisy predictions and completion of occluded boundaries using structural priors. A variational guidance network further regularizes initialization by encoding indoor layout priors. Panoramic benchmarks show that our diffusion approach provides a strong foundation for general room reconstruction, even in cluttered, non-Manhattan environments.

Reference and download information

Giovanni Pintore, Uzair Shah, Marco Agus, and Enrico Gobbetti. InLayDiffusion: Indoor Layout Estimation from a Single Panorama via Structural Point Cloud Lifting and Guided Diffusion. In Computer Graphics International. Volume 10605 of Lecture Notes in Computer Science (LNCS), Springer, 2026.

Related multimedia productions

Bibtex citation record

@incollection{Pintore:2026:IIL,
    author = {Giovanni Pintore and Uzair Shah and Marco Agus and Enrico Gobbetti},
    title = {{InLayDiffusion}: Indoor Layout Estimation from a Single Panorama via Structural Point Cloud Lifting and Guided Diffusion},
    booktitle = {Computer Graphics International},
    series = {Lecture Notes in Computer Science (LNCS)},
    volume = {10605},
    publisher = {Springer},
    year = {2026},
    abstract = { We introduce a novel end-to-end deep-learning method that combines structural geometric lifting with prior-guided diffusion to recover a 2D floorplan and a 3D room layout from a 360-degree image. Structural lifting predicts gravity-aligned, structure-aware depth and projects the resulting colored 3D point cloud onto the floor plane through a differentiable operation, directly linking panoramic image features to geometric reasoning and improving robustness to clutter while encoding cues such as room height and coarse footprint geometry. Footprint reconstruction is then formulated as polygon denoising and completion within a diffusion framework. The layout is represented as a fixed-vertex polygon jointly encoded with floor-projected RGB and density features, enabling refinement of noisy predictions and completion of occluded boundaries using structural priors. A variational guidance network further regularizes initialization by encoding indoor layout priors. Panoramic benchmarks show that our diffusion approach provides a strong foundation for general room reconstruction, even in cluttered, non-Manhattan environments. },
    url = {http://vic.crs4.it/vic/cgi-bin/bib-page.cgi?id='Pintore:2026:IIL'},
}