Novel unsupervised techniques for mixed-type data
File(s)
Author(s)
Costa, Efthymios
Type
Thesis
Abstract
Mixed-type data sets, that is, data sets consisting of continuous and categorical variables, are prevalent across multiple sectors, such as healthcare, finance, and social sciences. While many supervised methods account for this heterogeneity in the data, unsupervised techniques that rely on a Euclidean intuition are not directly applicable to non-continuous features. This underscores the need for unsupervised learning techniques which can handle categorical, as well as mixed-type data in a flexible manner. This thesis presents novel methodological developments in two fundamental tasks in unsupervised learning; cluster analysis and outlier detection.
The first part of the thesis explores existing non-model-based algorithms for clustering mixed-type data. It identifies the main strengths and weaknesses of specific methods and provides a taxonomy of eight state-of-the-art algorithms. Moreover, the most influential data set characteristics to the clustering performance are investigated, thus motivating the development of a novel clustering framework. The proposed framework leverages ideas from information theory, highlighting the versatility of this approach in integrating different variable types to the clustering problem.
In the second part of the thesis, the focus is on defining and detecting outliers in discrete and mixed-attribute domains. A definition of outlyingness for unordered categorical data is provided and a framework for outlier quantification is devised. The presented framework reflects on the considerations implied by the given definition and employs concepts from association rule mining to assign scores of anomalous behaviour to observations. The case of mixed-type data is also explored by extending ideas from the robust statistics literature for the case of mixed continuous-ordinal data. This provides a statistically principled approach to the problem of outlier identification. The extension to nominal data and the challenges associated with the latter are thoroughly discussed. Overall, this thesis introduces novel techniques for unsupervised learning problems in the presence of categorical and mixed-type data.
The first part of the thesis explores existing non-model-based algorithms for clustering mixed-type data. It identifies the main strengths and weaknesses of specific methods and provides a taxonomy of eight state-of-the-art algorithms. Moreover, the most influential data set characteristics to the clustering performance are investigated, thus motivating the development of a novel clustering framework. The proposed framework leverages ideas from information theory, highlighting the versatility of this approach in integrating different variable types to the clustering problem.
In the second part of the thesis, the focus is on defining and detecting outliers in discrete and mixed-attribute domains. A definition of outlyingness for unordered categorical data is provided and a framework for outlier quantification is devised. The presented framework reflects on the considerations implied by the given definition and employs concepts from association rule mining to assign scores of anomalous behaviour to observations. The case of mixed-type data is also explored by extending ideas from the robust statistics literature for the case of mixed continuous-ordinal data. This provides a statistically principled approach to the problem of outlier identification. The extension to nominal data and the challenges associated with the latter are thoroughly discussed. Overall, this thesis introduces novel techniques for unsupervised learning problems in the presence of categorical and mixed-type data.
Version
Open Access
Date Issued
2025-09-29
Date Awarded
2026-02-01
Copyright Statement
Attribution-Non Commercial-No Derivatives 4.0 International Licence (CC BY-NC-ND)
Advisor
Papatsouma, Ioanna
Young, Alastair
Publisher Department
Department of Mathematics
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
