Large-Scale Kernel Methods for Independence Testing
File(s)10.1007%2Fs11222-016-9721-7.pdf (934.96 KB)
Published version
Author(s)
Zhang, Q
Filippi, S
Gretton, A
Sejdinovic, D
Type
Journal Article
Abstract
Representations of probability measures in reproducing kernel Hilbert spaces
provide a flexible framework for fully nonparametric hypothesis tests of
independence, which can capture any type of departure from independence,
including nonlinear associations and multivariate interactions. However, these
approaches come with an at least quadratic computational cost in the number of
observations, which can be prohibitive in many applications. Arguably, it is
exactly in such large-scale datasets that capturing any type of dependence is
of interest, so striking a favourable tradeoff between computational efficiency
and test performance for kernel independence tests would have a direct impact
on their applicability in practice. In this contribution, we provide an
extensive study of the use of large-scale kernel approximations in the context
of independence testing, contrasting block-based, Nystrom and random Fourier
feature approaches. Through a variety of synthetic data experiments, it is
demonstrated that our novel large scale methods give comparable performance
with existing methods whilst using significantly less computation time and
memory.
provide a flexible framework for fully nonparametric hypothesis tests of
independence, which can capture any type of departure from independence,
including nonlinear associations and multivariate interactions. However, these
approaches come with an at least quadratic computational cost in the number of
observations, which can be prohibitive in many applications. Arguably, it is
exactly in such large-scale datasets that capturing any type of dependence is
of interest, so striking a favourable tradeoff between computational efficiency
and test performance for kernel independence tests would have a direct impact
on their applicability in practice. In this contribution, we provide an
extensive study of the use of large-scale kernel approximations in the context
of independence testing, contrasting block-based, Nystrom and random Fourier
feature approaches. Through a variety of synthetic data experiments, it is
demonstrated that our novel large scale methods give comparable performance
with existing methods whilst using significantly less computation time and
memory.
Date Issued
2017-01-24
Date Acceptance
2016-12-15
Citation
Statistics and Computing, 2017, 28 (1), pp.113-130
ISSN
1573-1375
Publisher
Springer Verlag (Germany)
Start Page
113
End Page
130
Journal / Book Title
Statistics and Computing
Volume
28
Issue
1
Copyright Statement
© The Author(s) 2017. This article is published with open access at Springerlink.com
License URL
Subjects
stat.CO
stat.CO
stat.ML
Notes
29 pages, 6 figures
Publication Status
Published