Semantic neural representation for SLAM and scene understanding
File(s)
Author(s)
Zhi, Shuaifeng
Type
Thesis
Abstract
Semantic simultaneous localisation and mapping (SLAM) has advanced remarkably over the past few years with the application of deep learning techniques. We are in the middle of a rapid progression from SLAM systems that simply reconstruct geometry towards understanding what is where in scenes. Semantically enriched maps will ultimately help intelligent robots improve the range and sophistication of their interactions with the world.
A key enabler of this capability is the scene representation, which defines how intelligent robots perceive, store and understand environmental attributes (i.e., physics, semantics, dynamics and interactions) from continuous observations. Geometric representation itself has long been a central research topic, and researchers have developed many ways to reconstruct the shape of scenes accurately and efficiently. We believe that representing geometry and semantics jointly is the right direction for scene models which are both optimally efficient and the most useful for actionable intelligence.
The work in this thesis concerns using tools from deep learning to enable joint representation of both geometry and semantics in SLAM systems and scene understanding. First we propose to learn code-based compact scene representations given large-scale datasets of images, depth maps and dense semantic labelling. The code representations of geometry and semantics are learned separately and enables joint optimisation at runtime, leading to a new multi-view label fusion approach and a preliminary dense monocular semantic mapping system. Secondly we explore a scene-specific implicit representation jointly encoding appearance, geometry and semantics. Without any external data, we show that the smoothness and consistency within our approach enable accurate semantic rendering given only sparse or noisy in-place annotation. Lastly, the real impact of this implicit representation in an super-efficient interactive scene labelling tool is shown.
A key enabler of this capability is the scene representation, which defines how intelligent robots perceive, store and understand environmental attributes (i.e., physics, semantics, dynamics and interactions) from continuous observations. Geometric representation itself has long been a central research topic, and researchers have developed many ways to reconstruct the shape of scenes accurately and efficiently. We believe that representing geometry and semantics jointly is the right direction for scene models which are both optimally efficient and the most useful for actionable intelligence.
The work in this thesis concerns using tools from deep learning to enable joint representation of both geometry and semantics in SLAM systems and scene understanding. First we propose to learn code-based compact scene representations given large-scale datasets of images, depth maps and dense semantic labelling. The code representations of geometry and semantics are learned separately and enables joint optimisation at runtime, leading to a new multi-view label fusion approach and a preliminary dense monocular semantic mapping system. Secondly we explore a scene-specific implicit representation jointly encoding appearance, geometry and semantics. Without any external data, we show that the smoothness and consistency within our approach enable accurate semantic rendering given only sparse or noisy in-place annotation. Lastly, the real impact of this implicit representation in an super-efficient interactive scene labelling tool is shown.
Version
Open Access
Date Issued
2021-09
Date Awarded
2021-12
Copyright Statement
Creative Commons Attribution-Non Commercial 4.0 International Licence
License URL
Advisor
Davison, Andrew
Leutenegger, Stefan
Sponsor
China Scholarship Council and Imperial College London
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)