Na Zhao
Singapore University of Technology and Design
Generalizable, Semantic, and Editable 3D Scene Reconstruction
Recent advances have made high-quality 3D reconstruction increasingly practical. However, photorealistic rendering alone is not enough for many downstream applications: reconstructed scenes should also generalize to unseen environments, capture meaningful semantic structure, and support intuitive editing and manipulation. In this talk, I will present our recent work toward this broader goal of generalizable, semantic, and editable 3D scene reconstruction. I will first discuss generalizable reconstruction from sparse views, focusing on how explicit geometric constraints and learned multi-view correspondence can be combined for robust and efficient reconstruction. I will then show how geometry can serve not only as an output, but also as a foundation for cross-view semantic reasoning, progressing from geometry–semantics synergy to open-vocabulary 3D semantic fields from unposed images. Finally, I will introduce compositional scene representations that combine interpretable geometric primitives with high-fidelity 3D Gaussians, enabling part-aware and physically meaningful editing. Together, these works explore how 3D reconstruction can evolve toward building structured, understandable, and actionable representations of the 3D world.