A multi-task framework for 2D scene parsing and 3D reconstruction in indoor space documentation

Erişen, Serdar and Mehranfar, Mansour and Borrmann, André; Moreno-Rangel, Alejandro and Kumar, Bimal, eds. (2025) A multi-task framework for 2D scene parsing and 3D reconstruction in indoor space documentation. In: EG-ICE 2025. University of Strathclyde Publishing, GBR, pp. 28-38. ISBN 9781914241826 (https://doi.org/10.17868/strath.00093243)

[thumbnail of Erisen-etal-EG-ICE-2025-A-multi-task-framework-for-2D-scene-parsing-and-3D-reconstruction]
Preview
Text. Filename: Erisen-etal-EG-ICE-2025-A-multi-task-framework-for-2D-scene-parsing-and-3D-reconstruction.pdf
Final Published Version
License: Creative Commons Attribution 4.0 logo

Download (1MB)| Preview

Abstract

Smart indoor environments with recognition technologies that prioritise human comfort and well-being necessitate accurate spatial information. Laser scanners and photogrammetry technologies have become essential for creating accurate as-built digital building models by capturing high-quality point clouds with precise geometry, yet they often lack proper semantic annotations. This limitation necessitates developing methods to improve 2D and 3D scene understanding. To address this, a multi-task learning framework is proposed combining depth estimation and semantic segmentation techniques to enhance 2D scene parsing for effective 3D reconstruction from single images by generating and learning 3D surface masks. The results on the selected TUM CMS Indoor Point Clouds dataset demonstrate the effectiveness of the proposed framework in 3D reconstruction, achieving 81.55% mIoU accuracy, which supports multiple applications for indoor digital twinning, automated 3D object recognition, and simulation tasks in indoor spaces.