Context-aware motion prediction of human-object interaction
File(s)
Author(s)
Razali, Haziq
Type
Thesis
Abstract
Understanding human motion is critical for the effective deployment of social and assistive robots that will interact with humans in shared environments. A key aspect of this understanding is the ability to predict and interpret human interactions with objects, which encompasses everyday tasks such as human-to-human object handovers, object retrieval and placement, and bimanual object manipulation. In this thesis, we present novel approaches to human and object motion forecasting, with a specific focus on scenarios involving hand-held objects.
We address the challenges posed by the complex and context-dependent nature of these tasks by developing methods that leverage various contextual cues, including eye gaze, object labels, shapes, and language. Our work spans several key areas: improving the accuracy of pose predictions during human-to-human object handovers, enhancing object retrieval and placement through the integration of gaze and language data, and advancing the \ac{SOTA} in bimanual object manipulation forecasting. Through the strategic use of contextual information and the application of modular network architectures, along with the introduction of keyframe forecasting to drive sequence generation, this thesis contributes to the development of more robust and natural motion forecasting systems.
The results demonstrate significant improvements in prediction errors, the ability to handle a diverse range of tasks, and enhanced performance in long-duration forecasting, laying the groundwork for future research in \ac{HRI} and related applications in \ac{AR} and \ac{VR}.
We address the challenges posed by the complex and context-dependent nature of these tasks by developing methods that leverage various contextual cues, including eye gaze, object labels, shapes, and language. Our work spans several key areas: improving the accuracy of pose predictions during human-to-human object handovers, enhancing object retrieval and placement through the integration of gaze and language data, and advancing the \ac{SOTA} in bimanual object manipulation forecasting. Through the strategic use of contextual information and the application of modular network architectures, along with the introduction of keyframe forecasting to drive sequence generation, this thesis contributes to the development of more robust and natural motion forecasting systems.
The results demonstrate significant improvements in prediction errors, the ability to handle a diverse range of tasks, and enhanced performance in long-duration forecasting, laying the groundwork for future research in \ac{HRI} and related applications in \ac{AR} and \ac{VR}.
Version
Open Access
Date Issued
2024-09-27
Date Awarded
2025-05-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Demiris, Yiannis
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
