On the efficiency of data collection for crowdsourced classification
File(s)0217.pdf (210.2 KB)
Published version
Author(s)
Manino, E
Tran-Thanh, L
Jennings, NR
Type
Conference Paper
Abstract
The quality of crowdsourced data is often highly variable. For this reason, it is common to collect redundant data and use statistical methods to aggregate it. Empirical studies show that the policies we use to collect such data have a strong impact on the accuracy of the system. However, there is little theoretical understanding of this phenomenon. In this paper we provide the first theoretical explanation of the accuracy gap between the most popular collection policies: the non-adaptive uniform allocation, and the adaptive uncertainty sampling and information gain maximisation. To do so, we propose a novel representation of the collection process in terms of random walks. Then, we use this tool to derive lower and upper bounds on the accuracy of the policies. With these bounds, we are able to quantify the advantage that the two adaptive policies have over the non-adaptive one for the first time.
Date Issued
2018-07-19
Date Acceptance
2018-07-13
Citation
IJCAI International Joint Conference on Artificial Intelligence, 2018, 2018, pp.1568-1575
ISBN
9780999241127
ISSN
1045-0823
Publisher
Lawrence Erlbaum Associates, Inc.
Start Page
1568
End Page
1575
Journal / Book Title
IJCAI International Joint Conference on Artificial Intelligence
Volume
2018
Copyright Statement
© 2018 International Joint Conferences on Artificial Intelligence. All rights reserved. No part of this book may be reproduced in any form by any electronic or mechanical means (including photocopying, recording, or information storage and retrieval) without permission in writing from the publisher.
Source
Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI-18)
Publication Status
Published
Start Date
2018-07-13
Finish Date
2018-07-19
Coverage Spatial
Stockholm, Sweden