SceneNet RGB-D: Can 5M synthetic images beat generic ImageNet pre-training on indoor segmentation?
File(s)scenenet_rgbd_post_iccv_submission.pdf (4.9 MB)
Accepted version
Author(s)
McCormac
Handa, A
Leutenegger, S
Davison, AJ
Type
Conference Paper
Abstract
We introduce SceneNet RGB-D, a dataset providing pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection. It also provides perfect camera poses and depth data, allowing investigation into geometric computer vision problems such as optical flow, camera pose estimation, and 3D scene labelling tasks. Random sampling permits virtually unlimited scene configurations, and here we provide 5M rendered RGB-D images from 16K randomly generated 3D trajectories in synthetic layouts, with random but physically simulated object configurations. We compare the semantic segmentation performance of network weights produced from pretraining on RGB images from our dataset against generic VGG-16 ImageNet weights. After fine-tuning on the SUN RGB-D and NYUv2 real-world datasets we find in both cases that the synthetically pre-trained network outperforms the VGG-16 weights. When synthetic pre-training includes a depth channel (something ImageNet cannot natively provide) the performance is greater still. This suggests that large-scale high-quality synthetic RGB datasets with task-specific labels can be more useful for pretraining than real-world generic pre-training such as ImageNet. We host the dataset at http://robotvault. bitbucket.io/scenenet-rgbd.html.
Date Issued
2017-12-25
Date Acceptance
2017-08-04
Citation
2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp.2697-2706
ISBN
9781538610329
ISSN
2380-7504
Publisher
IEEE
Start Page
2697
End Page
2706
Journal / Book Title
2017 IEEE International Conference on Computer Vision (ICCV)
Copyright Statement
© 2017 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
Dyson Technology Limited
Grant Number
PO 4500501004
Source
International Conference on Computer Vision 2017
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Engineering, Electrical & Electronic
Computer Science
Engineering
Publication Status
Published
Start Date
2017-10-22
Finish Date
2017-10-29
Coverage Spatial
Venice, Italy
Date Publish Online
2017-12-25