Distribution regression for sequential data
File(s)2006.05805v5.pdf (4.73 MB)
Published version
Author(s)
Lemercier, Maud
Salvi, Cristopher
Damoulas, Theodors
Bonilla, Edwin
Lyons, Terry
Type
Conference Paper
Abstract
Distribution regression refers to the supervised learning problem where labels are only available for groups of inputs instead of individual inputs. In this paper, we develop a rigorous mathematical framework for distribution regression where inputs are complex data streams. Leveraging properties of the expected signature and a recent signature kernel trick for sequential data from stochastic analysis, we introduce two new learning techniques, one feature-based and the other kernel-based. Each is suited to a different data regime in terms of the number of data streams and the dimensionality of the individual streams. We provide theoretical results on the universality of both approaches and demonstrate empirically their robustness to irregularly sampled multivariate time-series, achieving state-of-the-art performance on both synthetic and real-world examples from thermodynamics, mathematical finance and agricultural science.
Date Issued
2021-04-13
Date Acceptance
2021-01-22
Citation
Proceedings of Machine Learning Research, 2021, pp.3754-3762
ISSN
2640-3498
Publisher
MLResearchPress
Start Page
3754
End Page
3762
Journal / Book Title
Proceedings of Machine Learning Research
Copyright Statement
Copyright © The authors and PMLR 2023. MLResearchPress.
Identifier
http://proceedings.mlr.press/v130/lemercier21a.html
Source
The 24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021)
Publication Status
Published
Start Date
2021-04-13
Finish Date
2021-04-15
Coverage Spatial
Virtual