On-line adaptation of exploration in the one-armed bandit with covariates problem
OA Location
Author(s)
Sykulski, Adam M
Adams, Niall M
Jennings, Nicholas R
Type
Conference Paper
Abstract
Many sequential decision making problems require an agent to balance exploration and exploitation to maximise long-term reward. Existing policies that address this tradeoff typically have parameters that are set a priori to control the amount of exploration. In finite-time problems, the optimal values of these parameters are highly dependent on the problem faced. In this paper, we propose adapting the amount of exploration performed on-line, as information is gathered by the agent. To this end we introduce a novel algorithm, e-ADAPT, which has no free parameters. The algorithm adapts as it plays and sequentially chooses whether to explore or exploit, driven by the amount of uncertainty in the system. We provide simulation results for the onearmed bandit with covariates problem, which demonstrate the effectiveness of e-ADAPT to correctly control the amount of exploration in finite-time problems and yield rewards that are close to optimally tuned off-line policies. Furthermore, we show that e-ADAPT is robust to a high-dimensional covariate, as well as misspecified models. Finally, we describe how our methods could be extended to other sequential decision making problems, such as dynamic bandit problems. with changing reward structures.
Date Issued
2010-12
Date Acceptance
2011-02-04
Citation
2010, pp.459-464
Publisher
IEEE
Start Page
459
End Page
464
Journal / Book Title
Proceedings of the Internal Conference of Machine Learning and Applications, Washington DC, USA
Copyright Statement
© 2010 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Identifier
http://eprints.soton.ac.uk/271615/
Source
9th International Conference on Machine Learning and Applications (ICMLA)
Source Place
Washington, DC
Notes
Event Dates: 12-14 Dec, 2010 keywords: Exploration-exploitation tradeoff, sequential decision making, on-line learning, one-armed bandit problem
Publication Status
Unpublished
Start Date
2010-12-12
Finish Date
2010-12-14
