Accelerating 6D object pose estimation with scalable approaches
Author(s)
Vieira de Castro, Pedro
Type
Thesis
Abstract
For real-world applications, extracting 6D object pose of objects requires high accuracy, throughput and scalability. In this thesis, we propose solutions to address these needs.
Firstly, we design a GraphCNN based architecture specifically built for a single object. In contrast to previous works, we explicitly exploit the distinct topological information, 3D dense meshes of each object in the pose estimation model, with an automated process and prior to any post-processing refinement stage.
Previous methods overlook depth errors caused by the projective nature of the camera sensor. We propose Multi-Frustum Network, which uses a multi-frustum grid in allocentric space as an intermediate pose representation, simplifying the main task and enabling simultaneous detection and pose estimation within the same network for fast performance.
6D object pose projects typically train a network per object but optimizing a model for multiple objects sharply drops its accuracy. We propose applying object conditional parameterizations to a backbone network and distill knowledge from multiple pre-trained single-object teacher models to a multi-object student network.
6D object pose refinement methods rely on a slow render-compare pipeline. A rendering with a known object, camera and lighting represents its pose. We propose replacing the rendering step with a more efficient representation which is used in Cascaded Refinement Transformers. We achieve inference runtimes twice as fast as the nearest real-time state-of-the-art methods, supporting up to 30 objects in a single model without sacrificing accuracy.
Lastly, we introduce PoseMatcher, a one-shot object pose estimator that accurately estimates poses of unseen objects using few template images. At each training step, we emulate test-time scenarios by cheaply constructing an approximation of the full object point cloud with a three-view system: a query, positive and negative templates. Moreover, we redesign the transformer matching layer to accommodate different modalities and propose 3D based refinement module.
Firstly, we design a GraphCNN based architecture specifically built for a single object. In contrast to previous works, we explicitly exploit the distinct topological information, 3D dense meshes of each object in the pose estimation model, with an automated process and prior to any post-processing refinement stage.
Previous methods overlook depth errors caused by the projective nature of the camera sensor. We propose Multi-Frustum Network, which uses a multi-frustum grid in allocentric space as an intermediate pose representation, simplifying the main task and enabling simultaneous detection and pose estimation within the same network for fast performance.
6D object pose projects typically train a network per object but optimizing a model for multiple objects sharply drops its accuracy. We propose applying object conditional parameterizations to a backbone network and distill knowledge from multiple pre-trained single-object teacher models to a multi-object student network.
6D object pose refinement methods rely on a slow render-compare pipeline. A rendering with a known object, camera and lighting represents its pose. We propose replacing the rendering step with a more efficient representation which is used in Cascaded Refinement Transformers. We achieve inference runtimes twice as fast as the nearest real-time state-of-the-art methods, supporting up to 30 objects in a single model without sacrificing accuracy.
Lastly, we introduce PoseMatcher, a one-shot object pose estimator that accurately estimates poses of unseen objects using few template images. At each training step, we emulate test-time scenarios by cheaply constructing an approximation of the full object point cloud with a three-view system: a query, positive and negative templates. Moreover, we redesign the transformer matching layer to accommodate different modalities and propose 3D based refinement module.
Version
Open Access
Date Issued
2023-05-07
Date Awarded
01/04/2025
License URL
Advisor
Kim, Tae-Kyun
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
