Supervised machine learning via tensor networks
File(s)
Author(s)
Konstantinidis, Kriton
Type
Thesis
Abstract
Recent years have witnessed an unprecedented increase in the generation of manifold
types, volume and scale of data. Such a surge has rendered the standard tools of
linear algebra suboptimal in data processing and analysis, highlighting the need for
the development of more sophisticated mathematical algorithms that are able to
operate on exceedingly large and multi-dimensional data. A recently introduced
concept in Machine Learning (ML) that has the potential to fill this gap is Tensor
Networks (TNs), that is, networks of interconnected low-dimensional N-way arrays,
called tensors.
The increasing interest of the ML community in TNs stems from their suitability
to operate with both dense and sparse data, their Big Data compatible properties,
as well as their enhanced interpretability, owing to their multi-linear nature. The
applicability and potential of TNs have been demonstrated through state-of-the-art
(SOTA) performance in tasks including multi-modal learning, compression of large-
dimensional data, sequence to sequence learning, anomaly detection and theoretical
analysis of neural networks, to name but a few. Inspired by this potential, this thesis
sets out to develop novel supervised learning algorithms based on TNs, and explore
the numerous promising application avenues.
Initially focusing on deterministic models, a powerful TN framework for tabular
data (one of the most common data types where data are represented in a row-column
format) is introduced. Its versatility on both dense and sparse data, robustness to
hyperparameter choices, as well as enhanced interpretability are demonstrated over
both toy and real-life datasets. Next, the incorporation of graphs into TN models is
explored, through the development of a joint tensor-graph framework, the potential
of which is illustrated on a real-life financial experiment.
The focus is then directed towards probabilistic TN models, a research domain
that has received little attention in literature. Firstly, a TN approach to efficiently
learn Gaussian Process kernel embeddings is presented, which is shown to exhibit
enhanced interpretability while being competitive to SOTA neural network models.
Furthermore, a new family of models, Bayesian TNs (BTNs), is introduced; this is
the first work to extend TN models to purely probabilistic settings. BTNs are shown
to achieve comparable or better performance than Bayesian Neural Networks with
similar parameter complexity, while being able to provide credibility intervals for the
learnt parameters.
Finally, a hyperparameter optimisation framework based on TNs is proposed,
which outperforms several competitive SOTA approaches. It is the author’s hope
that the presented robust multi-linear models will help alleviate the interpretability
issue of current black box approaches while being able to provide comparable or
superior performance, both in the deterministic and probabilistic paradigms
types, volume and scale of data. Such a surge has rendered the standard tools of
linear algebra suboptimal in data processing and analysis, highlighting the need for
the development of more sophisticated mathematical algorithms that are able to
operate on exceedingly large and multi-dimensional data. A recently introduced
concept in Machine Learning (ML) that has the potential to fill this gap is Tensor
Networks (TNs), that is, networks of interconnected low-dimensional N-way arrays,
called tensors.
The increasing interest of the ML community in TNs stems from their suitability
to operate with both dense and sparse data, their Big Data compatible properties,
as well as their enhanced interpretability, owing to their multi-linear nature. The
applicability and potential of TNs have been demonstrated through state-of-the-art
(SOTA) performance in tasks including multi-modal learning, compression of large-
dimensional data, sequence to sequence learning, anomaly detection and theoretical
analysis of neural networks, to name but a few. Inspired by this potential, this thesis
sets out to develop novel supervised learning algorithms based on TNs, and explore
the numerous promising application avenues.
Initially focusing on deterministic models, a powerful TN framework for tabular
data (one of the most common data types where data are represented in a row-column
format) is introduced. Its versatility on both dense and sparse data, robustness to
hyperparameter choices, as well as enhanced interpretability are demonstrated over
both toy and real-life datasets. Next, the incorporation of graphs into TN models is
explored, through the development of a joint tensor-graph framework, the potential
of which is illustrated on a real-life financial experiment.
The focus is then directed towards probabilistic TN models, a research domain
that has received little attention in literature. Firstly, a TN approach to efficiently
learn Gaussian Process kernel embeddings is presented, which is shown to exhibit
enhanced interpretability while being competitive to SOTA neural network models.
Furthermore, a new family of models, Bayesian TNs (BTNs), is introduced; this is
the first work to extend TN models to purely probabilistic settings. BTNs are shown
to achieve comparable or better performance than Bayesian Neural Networks with
similar parameter complexity, while being able to provide credibility intervals for the
learnt parameters.
Finally, a hyperparameter optimisation framework based on TNs is proposed,
which outperforms several competitive SOTA approaches. It is the author’s hope
that the presented robust multi-linear models will help alleviate the interpretability
issue of current black box approaches while being able to provide comparable or
superior performance, both in the deterministic and probabilistic paradigms
Version
Open Access
Date Issued
2023-06
Date Awarded
2023-12
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Mandic, Danilo
Sponsor
Engineering and Physical Sciences Research Council
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)