Understanding the 3D spatial semantics of real-world environments is central to visual perception, scene reasoning, and 3D content generation. Real-world spaces are inherently three-dimensional, and robust 3D perception enables machines not only to comprehend spatial structures, object functions, and interactions, but also to generate, reconstruct, and edit plausible 3D scenes. Despite its importance, learning comprehensive 3D representations remains challenging due to limited 3D data, noisy or partial observations, and the high-dimensional and multimodal nature of the problem.
The primary objective of this project is to develop algorithmic approaches that advance 3D visual understanding and generation by learning robust, generalizable, and increasingly controllable representations of objects, scenes, and interactions. For increased generalizability, we exploit strong general priors learned from large-scale data in other modalities (e.g. text, 2D images, and video) and transfer them into 3D reasoning, particularly in settings where direct 3D supervision is scarce, such as interaction modeling and dynamic scene understanding. The research focuses on three complementary goals:
- Developing robust 3D semantic understanding under partial and noisy observations, enabling general reasoning about scene structure and relationships, while exploiting cross-modal priors to improve generalization.
- Introducing efficient, compact 3D-based representations and operators to encode and generate objects and scenes, leveraging intrinsic spatial properties to enable large-scale synthesis and reconstruction across diverse modalities and environments.
- Modeling functionality and interactions in 3D scenes, enabling spatio-temporal reasoning and instruction-driven manipulation of dynamic scenarios, with emphasis on transferring interaction knowledge from video and language data.
The expected impact potential of this work is substantial. By combining 3D understanding with generative modeling, the project will enable next-generation technologies in machine perception, immersive communications, mixed reality, architectural and industrial modeling, and robotics. It establishes a paradigm of spatially-grounded, 3D-consistent semantic understanding capable of both interpreting and generating complex environments, supporting machines that can perceive, reason about, and interact intelligently with the 3D world.