Alignist: CAD-informed orientation distribution estimation by fusing shape and correspondences
File(s)Aliginst_ECCV.pdf (2.2 MB)
Accepted version
Author(s)
Vutukur, Shishir Reddy
Haugaard, Rasmus Laurvig
Huang, Junwen
Busam, Benjamin
Birdal, Tolga
Type
Conference Paper
Abstract
Object pose distribution estimation is crucial in robotics for better path planning and handling of symmetric objects. Recent distribution estimation approaches employ contrastive learning-based approaches by maximizing the likelihood of a single pose estimate in the absence of a CAD model. We propose a pose distribution estimation method leveraging symmetry respecting correspondence distributions and shape information obtained using a CAD model. Contrastive learning-based approaches require an exhaustive amount of training images from differrent viewpoints to learn the distribution properly, which is not possible in realistic scenarios. Instead, we propose a pipeline that can leverage correspondence distributions and shape information from the CAD model, which are later used to learn pose distributions. Besides, having access
to pose distribution based on correspondences before learning pose distributions conditioned on images, can help formulate the loss between distributions. The prior knowledge of distribution also helps the network
to focus on getting sharper modes instead. With the CAD prior, our approach converges much faster and learns distribution better by focusing on learning sharper distribution near all the valid modes, unlike contrastive approaches, which focus on a single mode at a time. We achieve benchmark results on SYMSOL-I and T-Less datasets.
to pose distribution based on correspondences before learning pose distributions conditioned on images, can help formulate the loss between distributions. The prior knowledge of distribution also helps the network
to focus on getting sharper modes instead. With the CAD prior, our approach converges much faster and learns distribution better by focusing on learning sharper distribution near all the valid modes, unlike contrastive approaches, which focus on a single mode at a time. We achieve benchmark results on SYMSOL-I and T-Less datasets.
Date Issued
2024-10-20
Date Acceptance
2024-07-03
Citation
Lecture Notes in Computer Science, 2024, 15060, pp.351-369
ISBN
978-3-031-72627-9
Publisher
Springer Nature
Start Page
351
End Page
369
Journal / Book Title
Lecture Notes in Computer Science
Volume
15060
Copyright Statement
© 2025 The Author(s), under exclusive license to Springer Nature Switzerland AG. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Sponsor
Engineering & Physical Science Research Council (E
Grant Number
EP/X011364/1
Source
European Conference on Computer Vision (ECCV 2024)
Subjects
Artificial Intelligence & Image Processing
Publication Status
Published
Start Date
2024-09-29
Finish Date
2024-10-04
Coverage Spatial
Milan, Italy