Training and evaluating adversarial networks: from kernel discrepancies to applications
File(s)
Author(s)
Binkowski, Mikolaj
Type
Thesis
Abstract
Implicit generative modelling has experienced remarkable transformation with the emergence and expansion of Generative Adversarial Networks. This generic framework introduced by Goodfellow et al. [2014] became a leading paradigm in generative modelling of images, and allowed generation of sharp and convincing samples, often hardly distinguishable from the real pictures. On the other hand, this rapid growth created a range of problems, spawning a whole new area of research. Despite the impressive results, the adversarial min-max objective makes GANs notoriously difficult to train and requires various regularisation techniques. On the other hand, applications of GANs to non-visual domains have seen limited success. In this thesis we address problems coming from both of these lines of research.
Its first part focuses on the theoretical and empirical aspects of GAN formulations motivated by integral probability metrics. In particular, we study gradient penalty for MMD GANs, which use use Maximum Mean Discrepancy as the training objective, and clarify their relation to Wasserstein and Cramér GANs. The latter, being motivated by the energy distance, approximately matches the MMD with energy kernel, which we include it in the study of kernel choice for MMD GANs. We empirically show that gradient-penalised MMD GANs work well with smaller critic networks, improving the training time over Wasserstein GANs. We also present the properties of gradient estimators used in training of IPM-based GANs.
In the second part, we apply GANs to unsupervised image-to-image transfer and text-to-speech audio generation. Although impressive results have already been achieved in the former task, they remain limited to situations where modes of source and target distributions are well-matched. In addition, this area lacks rigorous probabilistic setup as well as clear metrics for evaluation. On the other hand, GANs for audio generation have so far worked only for relatively simple tasks. We address these problems by introducing batch weight for domain transfer and GAN-TTS, the first GAN able to generate high-fidelity speech from text. Both of these models benefit from changes to the discriminator networks: we apply joint discriminator on a product space of two domains to learn unsupervised transfer, and ensemble of random-window discriminators that operate on audio sub-samples of different lengths and frequencies.
Finally, we study the issue of GAN evaluation, critically examining the properties of the existing methods and proposing a range of new metrics for all of the aforementioned domains.
Its first part focuses on the theoretical and empirical aspects of GAN formulations motivated by integral probability metrics. In particular, we study gradient penalty for MMD GANs, which use use Maximum Mean Discrepancy as the training objective, and clarify their relation to Wasserstein and Cramér GANs. The latter, being motivated by the energy distance, approximately matches the MMD with energy kernel, which we include it in the study of kernel choice for MMD GANs. We empirically show that gradient-penalised MMD GANs work well with smaller critic networks, improving the training time over Wasserstein GANs. We also present the properties of gradient estimators used in training of IPM-based GANs.
In the second part, we apply GANs to unsupervised image-to-image transfer and text-to-speech audio generation. Although impressive results have already been achieved in the former task, they remain limited to situations where modes of source and target distributions are well-matched. In addition, this area lacks rigorous probabilistic setup as well as clear metrics for evaluation. On the other hand, GANs for audio generation have so far worked only for relatively simple tasks. We address these problems by introducing batch weight for domain transfer and GAN-TTS, the first GAN able to generate high-fidelity speech from text. Both of these models benefit from changes to the discriminator networks: we apply joint discriminator on a product space of two domains to learn unsupervised transfer, and ensemble of random-window discriminators that operate on audio sub-samples of different lengths and frequencies.
Finally, we study the issue of GAN evaluation, critically examining the properties of the existing methods and proposing a range of new metrics for all of the aforementioned domains.
Version
Open Access
Date Issued
2020-04
Date Awarded
2021-05
Copyright Statement
Creative Commons Attribution NonCommercial ShareAlike Licence
Advisor
Neuman, Eyal
Sponsor
Engineering and Physical Sciences Research Council
Publisher Department
Mathematics
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)