Learning and advancing local feature-based methods for 3D geometry systems
File(s)
Author(s)
Barroso Laguna, Axel
Type
Thesis
Abstract
This thesis focuses on finding the geometric relationship between two views. Learning about the geometry of a scene is a pivotal part of modern 3D geometry systems, and it is a popular and active area of research.
There are many methods to compute image geometry, and most of them rely on what we call image matching. Image matching aims at finding local correspondences between images, and although it involves multiple steps, image matching stages can be summarized into 1) finding characteristic local regions in the images, also known as keypoints, 2) calculating their mathematical representation, referred to as keypoint descriptors, and 3) searching the corresponding regions between the views. The first part of this thesis focuses on the detection of the keypoints. We study the similarities between handcrafted detectors and learned convolutional networks and propose a new architecture that combines handcrafted and learned filters, which achieves state-of-the-art results while being fast and lightweight. Then, we extend the idea of integrating handcrafted filters into the keypoint descriptor network and discuss its benefits in terms of matching capabilities. Moreover, we introduce a new objective function based on the discriminativeness of keypoint candidates, which guides the candidate keypoints toward easy-to-match regions and improves the robustness of our keypoint detector and descriptor. Finally, we study and isolate one geometric challenge in image matching, large scale changes between views. To deal with the scale challenge, we present a scale-aware matching pipeline that brings a solution to image pairs where previous multi-scale methods failed. Besides the ideas and
experiments discussed in this manuscript, all the presented methods and datasets
are publicly available on https://github.com/axelBarroso.
There are many methods to compute image geometry, and most of them rely on what we call image matching. Image matching aims at finding local correspondences between images, and although it involves multiple steps, image matching stages can be summarized into 1) finding characteristic local regions in the images, also known as keypoints, 2) calculating their mathematical representation, referred to as keypoint descriptors, and 3) searching the corresponding regions between the views. The first part of this thesis focuses on the detection of the keypoints. We study the similarities between handcrafted detectors and learned convolutional networks and propose a new architecture that combines handcrafted and learned filters, which achieves state-of-the-art results while being fast and lightweight. Then, we extend the idea of integrating handcrafted filters into the keypoint descriptor network and discuss its benefits in terms of matching capabilities. Moreover, we introduce a new objective function based on the discriminativeness of keypoint candidates, which guides the candidate keypoints toward easy-to-match regions and improves the robustness of our keypoint detector and descriptor. Finally, we study and isolate one geometric challenge in image matching, large scale changes between views. To deal with the scale challenge, we present a scale-aware matching pipeline that brings a solution to image pairs where previous multi-scale methods failed. Besides the ideas and
experiments discussed in this manuscript, all the presented methods and datasets
are publicly available on https://github.com/axelBarroso.
Version
Open Access
Date Issued
2022-04
Date Awarded
2022-12
Copyright Statement
Creative Commons Attribution Licence
License URL
Advisor
Mikolajczyk, Krystian
Sponsor
Department of Electrical and Electronic Engineering
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)