On agent incentives to manipulate human feedback in multi-agent reward learning scenarios
File(s)
Author(s)
Ward, Francis
Toni, francesca
Belardinelli, francesco
Type
Conference Paper
Abstract
In settings without well-defined goals, methods for reward learning
allow reinforcement learning agents to infer goals from human
feedback. Existing work has discussed the problem that such agents
may manipulate humans, or the reward learning process, in order
to gain higher reward. We introduce the neglected problem that, in
multi-agent settings, agents may have incentives to manipulate one
another’s reward functions in order to change each other’s behav-
ioral policies. We focus on the setting with humans acting alongside
assistive (artificial) agents who must learn the reward function by
interacting with these humans. We propose a possible solution to
manipulation of human feedback in this setting: the Shared Value
Prior (SVP). The SVP equips agents with an assumption that the
reward functions of all humans are similar. Given this assumption,
the actions of any human provide information to an agent about
its reward, and so the agent is incentivised to observe these actions
rather than to manipulate them. We present an expository example
in which the SVP prevents manipulation.
allow reinforcement learning agents to infer goals from human
feedback. Existing work has discussed the problem that such agents
may manipulate humans, or the reward learning process, in order
to gain higher reward. We introduce the neglected problem that, in
multi-agent settings, agents may have incentives to manipulate one
another’s reward functions in order to change each other’s behav-
ioral policies. We focus on the setting with humans acting alongside
assistive (artificial) agents who must learn the reward function by
interacting with these humans. We propose a possible solution to
manipulation of human feedback in this setting: the Shared Value
Prior (SVP). The SVP equips agents with an assumption that the
reward functions of all humans are similar. Given this assumption,
the actions of any human provide information to an agent about
its reward, and so the agent is incentivised to observe these actions
rather than to manipulate them. We present an expository example
in which the SVP prevents manipulation.
Date Issued
2022-05-09
Date Acceptance
2021-12-20
Citation
2022, pp.1759-1761
Publisher
ACM
Start Page
1759
End Page
1761
Copyright Statement
© 2022 International Foundation for Autonomous Agents and Multiagent
Systems (www.ifaamas.org). All rights reserved.
Systems (www.ifaamas.org). All rights reserved.
Identifier
https://ifaamas.org/Proceedings/aamas2022/
Source
AAMAS 22
Publication Status
Published
Start Date
2022-05-09
Finish Date
2022-05-13
Coverage Spatial
Virtual
Date Publish Online
2022-05-09
