Scalable 3D generative reconstruction for spatial AI
File(s)
Author(s)
Kong, Xin
Type
Thesis
Abstract
Spatial Artificial Intelligence (Spatial AI) aims to equip intelligent systems, such as robots, autonomous vehicles, and AR/VR devices, with the ability to perceive, map, and reason about the physical world in a structured and predictive manner. Achieving this requires 3D representations that are not only geometrically accurate but also scalable, compositional, and generative, enabling systems to reconstruct environments and synthesise consistent novel views. While classical SLAM systems and neural scene representations such as NeRF have achieved strong results in localisation and rendering, they remain limited by per-scene optimisation, limited scalability, and poor cross-scene generalisation. Meanwhile, recent advances in diffusion and autoregressive generative models have transformed image and video synthesis, yet extending these approaches to 3D and multi-view generation introduces new challenges in efficiency, consistency, and structured dependency modelling.
This thesis advances Spatial AI through three contributions addressing these challenges. First, we introduce vMAP, a vectorised object-level neural mapping framework that represents scenes as sets of learnable primitives. By combining object-centric decomposition with parallel optimisation, vMAP enables scalable and compositional neural mapping. Second, we propose EscherNet, a diffusion-based generative model for scalable multi-view synthesis. By learning cross-scene priors and decoupling generation from explicit 3D coordinates, EscherNet generalises to unseen environments and supports flexible conditioning on arbitrary input views. Finally, we present CausNVS, an autoregressive multi-view diffusion framework that introduces causal masking to model structured dependencies between reference and target views, enabling consistent and flexible view synthesis.
Together, these contributions advance Spatial AI toward scalable and generative 3D world modelling, demonstrating how structured mapping, diffusion-based synthesis, and autoregressive conditioning can be unified for predictive spatial understanding.
This thesis advances Spatial AI through three contributions addressing these challenges. First, we introduce vMAP, a vectorised object-level neural mapping framework that represents scenes as sets of learnable primitives. By combining object-centric decomposition with parallel optimisation, vMAP enables scalable and compositional neural mapping. Second, we propose EscherNet, a diffusion-based generative model for scalable multi-view synthesis. By learning cross-scene priors and decoupling generation from explicit 3D coordinates, EscherNet generalises to unseen environments and supports flexible conditioning on arbitrary input views. Finally, we present CausNVS, an autoregressive multi-view diffusion framework that introduces causal masking to model structured dependencies between reference and target views, enabling consistent and flexible view synthesis.
Together, these contributions advance Spatial AI toward scalable and generative 3D world modelling, demonstrating how structured mapping, diffusion-based synthesis, and autoregressive conditioning can be unified for predictive spatial understanding.
Version
Open Access
Date Issued
2025-10-10
Date Awarded
2026-04-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Davison, Andrew
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
