Learning classical planning strategies with policy gradient
File(s)main.pdf (404.66 KB)
Accepted version
Author(s)
Gomoluch, Pawel
Alrajeh, Dalal
Russo, Alessandra
Type
Conference Paper
Abstract
A common paradigm in classical planning is heuristic forward
search. Forward search planners often rely on simple
best-first search which remains fixed throughout the search
process. In this paper, we introduce a novel search framework
capable of alternating between several forward search
approaches while solving a particular planning problem. Selection
of the approach is performed using a trainable stochastic
policy, mapping the state of the search to a probability distribution
over the approaches. This enables using policy gradient
to learn search strategies tailored to a specific distributions
of planning problems and a selected performance metric,
e.g. the IPC score. We instantiate the framework by constructing
a policy space consisting of five search approaches
and a two-dimensional representation of the planner’s state.
Then, we train the system on randomly generated problems
from five IPC domains using three different performance metrics.
Our experimental results show that the learner is able
to discover domain-specific search strategies, improving the
planner’s performance relative to the baselines of plain bestfirst
search and a uniform policy.
search. Forward search planners often rely on simple
best-first search which remains fixed throughout the search
process. In this paper, we introduce a novel search framework
capable of alternating between several forward search
approaches while solving a particular planning problem. Selection
of the approach is performed using a trainable stochastic
policy, mapping the state of the search to a probability distribution
over the approaches. This enables using policy gradient
to learn search strategies tailored to a specific distributions
of planning problems and a selected performance metric,
e.g. the IPC score. We instantiate the framework by constructing
a policy space consisting of five search approaches
and a two-dimensional representation of the planner’s state.
Then, we train the system on randomly generated problems
from five IPC domains using three different performance metrics.
Our experimental results show that the learner is able
to discover domain-specific search strategies, improving the
planner’s performance relative to the baselines of plain bestfirst
search and a uniform policy.
Date Issued
2019-07
Date Acceptance
2019-02-07
Citation
Vol. 29 (2019): Proceedings of the Twenty-Ninth International Conference on Automated Planning and Scheduling, 2019, 29, pp.637-645
ISSN
2334-0843
Publisher
AAAI
Start Page
637
End Page
645
Journal / Book Title
Vol. 29 (2019): Proceedings of the Twenty-Ninth International Conference on Automated Planning and Scheduling
Volume
29
Copyright Statement
© 2019, Association for the Advancement of Artificial Intelligence. All rights reserved.
Source
International Conference on Automated Planning and Scheduling
Publication Status
Published
Start Date
2019-07-11
Finish Date
2019-07-15
Coverage Spatial
Berkeley, CA, USA
Date Publish Online
2019-07-06