Advancements in face reconstruction from a single image
File(s)
Author(s)
Galanakis, Efstathios
Type
Thesis
Abstract
This Thesis addresses the fundamental challenge of 3D face reconstruction from a single image. Since the introduction of the seminal 3D Morphable Model (3DMM), significant advancements have been made in both human face representations, including implicit functions, and generation approaches, such as GANs and diffusion models. We explore the trade-off between these representations and generation techniques by proposing approaches that effectively leverage their strengths of both, for achieving high-fidelity results.
Firstly, we present 3DMM-RF, an implicit 3DMM that integrates a neural radiance field to a style-based GAN for face generation. By synthesizing the radiance field in one-pass, we overcome NeRF’s limitations of rigidity and slow rendering. Whilst trained on a large synthetic dataset, our model accurately reconstructs facial identities, under arbitrary pose, and appearance, enabling controllable re-rendering of “in-the-wild” images.
Next, we introduce FitDiff, a multi-modal latent diffusion model that generates relightable 3D facial avatars. It jointly outputs facial reflectance UV maps (diffuse albedo, specular albedo and normals) and shape given an input 2D image. During sampling, the proposed methodology achieves state-of-the-art 3D facial fitting as it utilizes a robust identity conditioning mechanism while employing perceptual and identity losses that guide the process.
Finally, SpinMeRound is a diffusion-based approach that operates directly in the image space, which can be used for synthesizing consistent all-around multi-view head portraits. Given an input facial image, it generates novel viewpoints of it, while preserving the crucial identity features. Our experiments show that SpinMeRound surpasses existing multi-view diffusion models in full-head synthesis.
Firstly, we present 3DMM-RF, an implicit 3DMM that integrates a neural radiance field to a style-based GAN for face generation. By synthesizing the radiance field in one-pass, we overcome NeRF’s limitations of rigidity and slow rendering. Whilst trained on a large synthetic dataset, our model accurately reconstructs facial identities, under arbitrary pose, and appearance, enabling controllable re-rendering of “in-the-wild” images.
Next, we introduce FitDiff, a multi-modal latent diffusion model that generates relightable 3D facial avatars. It jointly outputs facial reflectance UV maps (diffuse albedo, specular albedo and normals) and shape given an input 2D image. During sampling, the proposed methodology achieves state-of-the-art 3D facial fitting as it utilizes a robust identity conditioning mechanism while employing perceptual and identity losses that guide the process.
Finally, SpinMeRound is a diffusion-based approach that operates directly in the image space, which can be used for synthesizing consistent all-around multi-view head portraits. Given an input facial image, it generates novel viewpoints of it, while preserving the crucial identity features. Our experiments show that SpinMeRound surpasses existing multi-view diffusion models in full-head synthesis.
Version
Open Access
Date Issued
2025-04-04
Date Awarded
01/01/2026
License URL
Advisor
Zafeiriou, Stefanos
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
