Real-world human-robot collaborative reinforcement learning
File(s) 2003.01156v1.pdf (5.07 MB)
Working paper
Author(s)
Shafti, Ali
Tjomsland, Jonas
Dudley, William
Faisal, A Aldo
Type
Working Paper
Abstract
The intuitive collaboration of humans and intelligent robots (embodied AI) in
the real-world is an essential objective for many desirable applications of
robotics. Whilst there is much research regarding explicit communication, we
focus on how humans and robots interact implicitly, on motor adaptation level.
We present a real-world setup of a human-robot collaborative maze game,
designed to be non-trivial and only solvable through collaboration, by limiting
the actions to rotations of two orthogonal axes, and assigning each axes to one
player. This results in neither the human nor the agent being able to solve the
game on their own. We use a state-of-the-art reinforcement learning algorithm
for the robotic agent, and achieve results within 30 minutes of real-world
play, without any type of pre-training. We then use this system to perform
systematic experiments on human/agent behaviour and adaptation when co-learning
a policy for the collaborative game. We present results on how co-policy
learning occurs over time between the human and the robotic agent resulting in
each participant's agent serving as a representation of how they would play the
game. This allows us to relate a person's success when playing with different
agents than their own, by comparing the policy of the agent with that of their
own agent.
the real-world is an essential objective for many desirable applications of
robotics. Whilst there is much research regarding explicit communication, we
focus on how humans and robots interact implicitly, on motor adaptation level.
We present a real-world setup of a human-robot collaborative maze game,
designed to be non-trivial and only solvable through collaboration, by limiting
the actions to rotations of two orthogonal axes, and assigning each axes to one
player. This results in neither the human nor the agent being able to solve the
game on their own. We use a state-of-the-art reinforcement learning algorithm
for the robotic agent, and achieve results within 30 minutes of real-world
play, without any type of pre-training. We then use this system to perform
systematic experiments on human/agent behaviour and adaptation when co-learning
a policy for the collaborative game. We present results on how co-policy
learning occurs over time between the human and the robotic agent resulting in
each participant's agent serving as a representation of how they would play the
game. This allows us to relate a person's success when playing with different
agents than their own, by comparing the policy of the agent with that of their
own agent.
Date Issued
2020-03-02
Citation
2020
Publisher
arXiv
Copyright Statement
© 2020 The Author(s)
Identifier
http://arxiv.org/abs/2003.01156v1
Subjects
cs.RO
cs.RO
cs.AI
cs.LG
Notes
6 pages - under review at IROS2020
Publication Status
Published
