Pairwise Decomposition of Image Sequences for Active Multi-View Recognition
File(s)cvpr_2016_ed_johns.pdf (968.03 KB)
Accepted version
Author(s)
Johns, E
Leutenegger, S
Davison, AJ
Type
Conference Paper
Abstract
A multi-view image sequence provides a much richer
capacity for object recognition than from a single image.
However, most existing solutions to multi-view recognition
typically adopt hand-crafted, model-based geometric methods,
which do not readily embrace recent trends in deep
learning. We propose to bring Convolutional Neural Networks
to generic multi-view recognition, by decomposing
an image sequence into a set of image pairs, classifying
each pair independently, and then learning an object classi-
fier by weighting the contribution of each pair. This allows
for recognition over arbitrary camera trajectories, without
requiring explicit training over the potentially infinite number
of camera paths and lengths. Building these pairwise
relationships then naturally extends to the next-best-view
problem in an active recognition framework. To achieve
this, we train a second Convolutional Neural Network to
map directly from an observed image to next viewpoint.
Finally, we incorporate this into a trajectory optimisation
task, whereby the best recognition confidence is sought for
a given trajectory length. We present state-of-the-art results
in both guided and unguided multi-view recognition on the
ModelNet dataset, and show how our method can be used
with depth images, greyscale images, or both.
capacity for object recognition than from a single image.
However, most existing solutions to multi-view recognition
typically adopt hand-crafted, model-based geometric methods,
which do not readily embrace recent trends in deep
learning. We propose to bring Convolutional Neural Networks
to generic multi-view recognition, by decomposing
an image sequence into a set of image pairs, classifying
each pair independently, and then learning an object classi-
fier by weighting the contribution of each pair. This allows
for recognition over arbitrary camera trajectories, without
requiring explicit training over the potentially infinite number
of camera paths and lengths. Building these pairwise
relationships then naturally extends to the next-best-view
problem in an active recognition framework. To achieve
this, we train a second Convolutional Neural Network to
map directly from an observed image to next viewpoint.
Finally, we incorporate this into a trajectory optimisation
task, whereby the best recognition confidence is sought for
a given trajectory length. We present state-of-the-art results
in both guided and unguided multi-view recognition on the
ModelNet dataset, and show how our method can be used
with depth images, greyscale images, or both.
Date Issued
2016-12-12
Date Acceptance
2016-04-11
Citation
2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
ISSN
1063-6919
Publisher
Computer Vision Foundation (CVF)
Journal / Book Title
2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Copyright Statement
© 2016 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
Dyson Technology Limited
Grant Number
PO 4500378543
Source
Computer Vision and Pattern Recognition
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science
cs.CV
cs.RO
Publication Status
Published
Start Date
2016-06-26
Finish Date
2016-07-01
Coverage Spatial
Las Vegas, Nevada USA