Unsupervised learning of human movement from images
File(s)
Author(s)
Schmidtke, Luca
Type
Thesis
Abstract
Capturing and understanding human movement through machines has many applications including animation, virtual- and augmented reality, sports as well as medicine. All these benefit from the advent of increasingly powerful, data-driven artificial neural networks with millions or billions of parameters. This advancement, however, is accompanied by an ever-surging demand for more computational resources.
Data with which intelligent systems can be trained comes in two forms: with or without annotations. In the supervised training paradigm, models can be trained to predict these annotations coming from human input. Data without annotations however is many times more abundant. Humans have the remarkable ability to learn an internal world-model, i.e. a rich set of rules that govern the world in which they live in from different sensory inputs without such supervision. A central goal of artificial intelligence research is to enable machines to do something similar.
The central topic of this thesis is to investigate how to capture human movement from images through self-supervision. This movement can be characterised by pose, a set of points that denote the location of anatomical landmarks such as skeletal joints. The central theme throughout all the chapters is learning a disentangled representation of pose and appearance of any person in an image though a computational model.
In this thesis we find that a human shape template, a collection of connected, basic shapes, that serves as an abstract representation for pose, provides a strong prior for unsupervised pose estimation. Different self-supervision mechanisms such as conditional image translation as well as generation are introduced in combination with this shape template. We find that we can learn human pose and appearance from images in a flexible way and with minimal supervision, paving the way towards machine learning systems that can leverage unprecedented amounts of data.
Data with which intelligent systems can be trained comes in two forms: with or without annotations. In the supervised training paradigm, models can be trained to predict these annotations coming from human input. Data without annotations however is many times more abundant. Humans have the remarkable ability to learn an internal world-model, i.e. a rich set of rules that govern the world in which they live in from different sensory inputs without such supervision. A central goal of artificial intelligence research is to enable machines to do something similar.
The central topic of this thesis is to investigate how to capture human movement from images through self-supervision. This movement can be characterised by pose, a set of points that denote the location of anatomical landmarks such as skeletal joints. The central theme throughout all the chapters is learning a disentangled representation of pose and appearance of any person in an image though a computational model.
In this thesis we find that a human shape template, a collection of connected, basic shapes, that serves as an abstract representation for pose, provides a strong prior for unsupervised pose estimation. Different self-supervision mechanisms such as conditional image translation as well as generation are introduced in combination with this shape template. We find that we can learn human pose and appearance from images in a flexible way and with minimal supervision, paving the way towards machine learning systems that can leverage unprecedented amounts of data.
Version
Open Access
Date Issued
2024-03-25
Date Awarded
01/02/2025
License URL
Advisor
Kainz, Bernhard
Sponsor
Engineering and Physical Sciences Research Council
Grant Number
EP/S013687/1
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
