Orthogonally decoupled variational Gaussian processes
File(s)1809.08820v1.pdf (940.79 KB)
Accepted version
Author(s)
Salimbeni, HR
Cheng, Ching-An
Boots, Byron
Deisenroth, MP
Type
Conference Paper
Abstract
Gaussian processes (GPs) provide a powerful non-parametric framework for rea-
soning over functions. Despite appealing theory, its superlinear computational and
memory complexities have presented a long-standing challenge. State-of-the-art
sparse variational inference methods trade modeling accuracy against complexity.
However, the complexities of these methods still scale superlinearly in the number
of basis functions, implying that that sparse GP methods are able to learn from
large datasets only when a small model is used. Recently, a decoupled approach
was proposed that removes the unnecessary coupling between the complexities
of modeling the mean and the covariance functions of a GP. It achieves a linear
complexity in the number of mean parameters, so an expressive posterior mean
function can be modeled. While promising, this approach suffers from optimization
difficulties due to ill-conditioning and non-convexity. In this work, we propose an
alternative decoupled parametrization. It adopts an orthogonal basis in the mean
function to model the residues that cannot be learned by the standard coupled ap-
proach. Therefore, our method extends, rather than replaces, the coupled approach
to achieve strictly better performance. This construction admits a straightforward
natural gradient update rule, so the structure of the information manifold that is
lost during decoupling can be leveraged to speed up learning. Empirically, our
algorithm demonstrates significantly faster convergence in multiple experiments.
soning over functions. Despite appealing theory, its superlinear computational and
memory complexities have presented a long-standing challenge. State-of-the-art
sparse variational inference methods trade modeling accuracy against complexity.
However, the complexities of these methods still scale superlinearly in the number
of basis functions, implying that that sparse GP methods are able to learn from
large datasets only when a small model is used. Recently, a decoupled approach
was proposed that removes the unnecessary coupling between the complexities
of modeling the mean and the covariance functions of a GP. It achieves a linear
complexity in the number of mean parameters, so an expressive posterior mean
function can be modeled. While promising, this approach suffers from optimization
difficulties due to ill-conditioning and non-convexity. In this work, we propose an
alternative decoupled parametrization. It adopts an orthogonal basis in the mean
function to model the residues that cannot be learned by the standard coupled ap-
proach. Therefore, our method extends, rather than replaces, the coupled approach
to achieve strictly better performance. This construction admits a straightforward
natural gradient update rule, so the structure of the information manifold that is
lost during decoupling can be leveraged to speed up learning. Empirically, our
algorithm demonstrates significantly faster convergence in multiple experiments.
Date Issued
2018-12-31
Date Acceptance
2018-09-05
Citation
Advances in Neural Information Processing Systems
ISSN
1049-5258
Publisher
Massachusetts Institute of Technology Press
Journal / Book Title
Advances in Neural Information Processing Systems
Copyright Statement
¨ 2018 The Author(s)
Source
Advances in Neural Information Processing Systems (NIPS) 2018
Subjects
stat.ML
cs.LG
1701 Psychology
1702 Cognitive Science
Publication Status
Accepted
Start Date
2018-12-02
Finish Date
2018-12-08
Coverage Spatial
Montreal, Canada