Improving monocular 3D object detection by modeling geometry and uncertainty
File(s)
Author(s)
Shi, Xuepeng
Type
Thesis
Abstract
Monocular 3D object detection is the task of detecting objects as 3D bounding boxes in the physical space, using monocular images as input. It is significant for autonomous driving, robotics, and augmented reality because of the high availability of monocular cameras. However, it remains challenging because of the large distance variations of objects, the lack of explicit depth information from monocular images, and the inherent ambiguity in extracting 3D information from monocular images. Therefore, geometric priors, such as the physical size of objects and the imaging processing of cameras, and uncertainty modeling are crucial for monocular 3D object detection.
To handle the challenges, this thesis utilizes geometric considerations and models uncertainties in monocular 3D object detection for urban autonomous driving. Firstly, the large distance variations of objects are tackled by introducing a distance normalization method to learn a unified representation of objects in different scale and distance ranges. This framework can utilize the network capacity more efficiently.
Secondly, a geometry-based distance decomposition method is proposed to achieve accurate and interpretable distance prediction. This method applies explicit constraints on the distance prediction of objects by recovering the distance by its factors, unlike regressing it as a single variable in most existing methods. Uncertainty modeling of the factors is also introduced, considering the inherent ambiguity in recovering 3D information from monocular images.
Thirdly, a multivariate probabilistic modeling approach is proposed to improve the geometry-based distance decomposition method. In this framework, the joint probability distribution of the distance factors is explicitly modeled by learning the full covariance matrix of the factors.
The proposed methodologies can achieve state-of-the-art performance on monocular 3D object detection benchmarks, demonstrating their potential to advance the boundary and enable applications in autonomous driving, robotics, augmented reality, and beyond.
To handle the challenges, this thesis utilizes geometric considerations and models uncertainties in monocular 3D object detection for urban autonomous driving. Firstly, the large distance variations of objects are tackled by introducing a distance normalization method to learn a unified representation of objects in different scale and distance ranges. This framework can utilize the network capacity more efficiently.
Secondly, a geometry-based distance decomposition method is proposed to achieve accurate and interpretable distance prediction. This method applies explicit constraints on the distance prediction of objects by recovering the distance by its factors, unlike regressing it as a single variable in most existing methods. Uncertainty modeling of the factors is also introduced, considering the inherent ambiguity in recovering 3D information from monocular images.
Thirdly, a multivariate probabilistic modeling approach is proposed to improve the geometry-based distance decomposition method. In this framework, the joint probability distribution of the distance factors is explicitly modeled by learning the full covariance matrix of the factors.
The proposed methodologies can achieve state-of-the-art performance on monocular 3D object detection benchmarks, demonstrating their potential to advance the boundary and enable applications in autonomous driving, robotics, augmented reality, and beyond.
Version
Open Access
Date Issued
2023-07
Date Awarded
2024-06
Copyright Statement
Creative Commons Attribution NonCommercial NoDerivatives Licence
Advisor
Kim, Tae-Kyun
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)