H-Net: unsupervised attention-based stereo depth estimation leveraging epipolar geometry
Author(s)
Huang, Baoru
Zheng, Jian-Qing
Giannarou, Stamatia
Elson, Daniel S
Type
Conference Paper
Abstract
Depth estimation from a stereo image pair has become one of the most explored applications in computer vision, with most previous methods relying on fully supervised learning settings. However, due to the difficulty in acquiring accurate and scalable ground truth data, the training of fully supervised methods is challenging. As an alternative, self-supervised methods are becoming more popular to mitigate this challenge. In this paper, we introduce the H-Net, a deep-learning framework for unsupervised stereo depth estimation that leverages epipolar geometry to refine stereo matching. For the first time, a Siamese autoencoder architecture is used for depth estimation which allows mutual information between rectified stereo images to be extracted. To enforce the epipolar constraint, the mutual epipolar attention mechanism has been designed which gives more emphasis to correspondences of features that lie on the same epipolar line while learning mutual information between the input stereo pair. Stereo correspondences are further enhanced by incorporating semantic information to the proposed attention mechanism. More specifically, the optimal transport algorithm is used to suppress attention and eliminate outliers in areas not visible in both cameras. Extensive experiments on KITTI2015 and Cityscapes show that the proposed modules are able to improve the performance of the unsupervised stereo depth estimation methods while closing the gap with the fully supervised approaches.
Date Issued
2022-08-23
Date Acceptance
2022-08-01
Citation
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS (CVPRW 2022), 2022, pp.4459-4466
Publisher
IEEE
Start Page
4459
End Page
4466
Journal / Book Title
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS (CVPRW 2022)
Copyright Statement
Copyright © 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
National Institute for Health Research
Cancer Research UK
Imperial College Healthcare NHS Trust- BRC Funding
Identifier
https://www.webofscience.com/api/gateway?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000861612704057&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=a2bf6146997ec60c407a63945d4e92bb
Grant Number
NIHR200035
25147
RDB04
Source
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science, Theory & Methods
Computer Science
Publication Status
Published
Start Date
2022-06-18
Finish Date
2022-06-24
Coverage Spatial
New Orleans, LA
Date Publish Online
2022-08-23
