Differentiable rendering for dense visual SLAM
File(s)
Author(s)
Matsuki, Hidenobu
Type
Thesis
Abstract
This thesis investigates differentiable rendering for dense visual Simultaneous Localization and Mapping (SLAM), specifically examining how novel rendering algorithms and 3D scene representations enhance reconstruction and spatial understanding. The choice of scene representation is fundamental to the efficiency of 3D modeling. Traditional approaches, which rely on direct 3D measurements from sensors or predictions from pre-trained networks, often struggle with limited operational ranges or training data biases. This necessitates 3D representations that can be flexibly optimized using real-time 2D observations.
With advancements in computing hardware, 3D representations optimized through differentiable rendering have gained prominence. These methods achieve high-fidelity inverse rendering and novel view synthesis. While early techniques required lengthy optimization, the introduction of explicit data structures has shifted performance toward real-time, enabling integration into SLAM. Unlike offline reconstruction, online SLAM requires continuous estimation and updating of unknown environments from sequential data, posing unique challenges for optimization stability and memory management.
This thesis examines the use of 3D scene representations in combination with differentiable rendering pipelines to enable real-time SLAM systems. The work begins by developing classical depth map based SLAM methods to establish a foundation in existing standard pipelines. Building on this baseline, the thesis explores advanced scene representations — specifically NeRF and the latest 3D Gaussian Splatting techniques — as novel approaches to SLAM. Through detailed experiments and analysis, the thesis demonstrates how to effectively address the open-world challenges inherent in SLAM tasks using these modern representations. It also highlights the unique advantages that differentiable rendering offers, such as improved accuracy and flexibility, positioning it as a promising direction for the future of visual SLAM.
With advancements in computing hardware, 3D representations optimized through differentiable rendering have gained prominence. These methods achieve high-fidelity inverse rendering and novel view synthesis. While early techniques required lengthy optimization, the introduction of explicit data structures has shifted performance toward real-time, enabling integration into SLAM. Unlike offline reconstruction, online SLAM requires continuous estimation and updating of unknown environments from sequential data, posing unique challenges for optimization stability and memory management.
This thesis examines the use of 3D scene representations in combination with differentiable rendering pipelines to enable real-time SLAM systems. The work begins by developing classical depth map based SLAM methods to establish a foundation in existing standard pipelines. Building on this baseline, the thesis explores advanced scene representations — specifically NeRF and the latest 3D Gaussian Splatting techniques — as novel approaches to SLAM. Through detailed experiments and analysis, the thesis demonstrates how to effectively address the open-world challenges inherent in SLAM tasks using these modern representations. It also highlights the unique advantages that differentiable rendering offers, such as improved accuracy and flexibility, positioning it as a promising direction for the future of visual SLAM.
Version
Open Access
Date Issued
2024-12-31
Date Awarded
01/02/2026
Advisor
Davison, Andrew
Sponsor
James Dyson Foundation
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
