Multi-object scene completion using view-based network predictions
File(s)
Author(s)
Nicastro, Andrea
Type
Thesis
Abstract
Recent years have witnessed a commoditisation of autonomous robots. Almost-self-driving cars ride the roads and vacuum cleaners roam freely in houses. While these advancements showcase the navigation of an agent within the world, more complex interactions are yet to be unlocked.
More complex tasks, such as manipulation, require an additional layer of abstraction of how the world is represented. An agent is not trying to navigate an environment, any more. Rather, the agent interacts with entities and tries to control their state. The nature of the agents depends on the task. For example, manipulation tasks involve handling objects.
The state of the agent and a representation of the world are the basic building block of a task execution pipeline. Visual Simultaneous Localisation and Mapping (Visual SLAM) coupled with three-dimensional reconstruction algorithms have proven effective in building a dense representation of the geometry of the environment. However, reconstructing objects and scenes can take time and can delay the execution of actions.
Hence, this thesis elaborates on the concept of reconstructing objects and, specifically, on how to recover the complete shape of objects from incomplete geometry. While different shape representations exist, such as voxels and meshes, this document explores the use of two-dimensional representations of geometry, such as depths and thickness images. Moreover, here it is suggested a multiple-stage pipeline to break a scene into entities before recovering missing information.
More complex tasks, such as manipulation, require an additional layer of abstraction of how the world is represented. An agent is not trying to navigate an environment, any more. Rather, the agent interacts with entities and tries to control their state. The nature of the agents depends on the task. For example, manipulation tasks involve handling objects.
The state of the agent and a representation of the world are the basic building block of a task execution pipeline. Visual Simultaneous Localisation and Mapping (Visual SLAM) coupled with three-dimensional reconstruction algorithms have proven effective in building a dense representation of the geometry of the environment. However, reconstructing objects and scenes can take time and can delay the execution of actions.
Hence, this thesis elaborates on the concept of reconstructing objects and, specifically, on how to recover the complete shape of objects from incomplete geometry. While different shape representations exist, such as voxels and meshes, this document explores the use of two-dimensional representations of geometry, such as depths and thickness images. Moreover, here it is suggested a multiple-stage pipeline to break a scene into entities before recovering missing information.
Version
Open Access
Date Issued
2021-05
Date Awarded
2021-10
Copyright Statement
Creative Commons Attribution NonCommercial Licence
Advisor
Leutenegger, Stefan
Davison, Andrew
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)