Sample-Based Policy Iteration for Constrained DEC-POMDPs
File(s)FAIA242-0858.pdf (244.43 KB)
Published version
OA Location
Author(s)
Wu, F
Jennings, N
Chen, X
Type
Conference Paper
Abstract
We introduce constrained DEC-POMDPs – an extension of the standard DEC-POMDPs that includes constraints on the optimality of the overall team rewards. Constrained DEC-POMDPs present a natural framework for modeling cooperative multi-agent problems with limited resources. To solve such DEC-POMDPs, we propose a novel sample-based policy iteration algorithm. The algorithm builds on multi-agent dynamic programming and benefits from several recent advances in DEC-POMDP algorithms such as MBDP [12] and TBDP [13]. Specifically, it improves the joint policy by solving a series of standard nonlinear programs (NLPs), thereby building on recent advances in NLP solvers. Our experimental results confirm the algorithm can efficiently solve constrained DECPOMDPs that cause general DEC-POMDP algorithms to fail.
Date Issued
2012-08-27
Date Acceptance
2012-08-27
Citation
Frontiers in Artificial Intelligence and Applications, 2012, 242, pp.858-863
ISSN
1535-6698
Publisher
IOS Press
Start Page
858
End Page
863
Journal / Book Title
Frontiers in Artificial Intelligence and Applications
Volume
242
Copyright Statement
© 2012 The Author(s).
This article is published online with Open Access by IOS Press and distributed under the terms
of the Creative Commons Attribution Non-Commercial License
This article is published online with Open Access by IOS Press and distributed under the terms
of the Creative Commons Attribution Non-Commercial License
License URL
Identifier
http://eprints.soton.ac.uk/339937/
Source
20th European Conference on Artificial Intelligence (ECAI 2012)
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science
Publication Status
Unpublished
Start Date
2012-08-27
Finish Date
2012-08-31
Coverage Spatial
Montpellier, France