Mixed data regression for price prediction
File(s)
Author(s)
Larionov, Alexander
Type
Thesis
Abstract
There are many applications of price regression: to manage cash flows for business, balance supply and demand for energy, or quantify risks for insurance.
Historically, regression models -- be they linear or complex neural networks -- have mainly used `tabular' data composed of numeric or nominal categorical variables.
However, as companies and institutes increasingly collect more text data, and language models substantially improve every year, there is growing interest in how to leverage text for regression.
This thesis studies using text-only data as well as combining text and tabular data for regression, applied to price prediction.
We begin by discussing an often overlooked aspect of tabular modelling: missing value handling.
A large, real UK insurer dataset is used to study the impact of missing value handling on the performance of gradient boosted trees and deep learning models, two contemporary tabular modelling approaches.
In doing so we establish Catboost as a strong tabular model, relevant for benchmarking joint text-tabular models, but emphasize the importance of missing values.
We then conduct a study of text-only regression.
Text-only regression requires a pipeline of preprocessing, tokenization, featurization and modelling steps.
We rigorously analyze choices for these steps, where the literature lacks systematic work.
Among other results, for the first time, our study quantifies the relative importance of the model choice compared to choice of steps before modelling.
We find that choice of the upstream steps contributes variance in performance comparable to model choice.
We conclude with a study of joint text-tabular modelling: studying an ensemble of a tabular-only Catboost and a text-only neural network.
This ensemble is compared to existing methods.
Overall, we conclude that although this ensemble does not consistently have the best accuracy, it is simple, robust to data with low signal in the text modality and achieves good accuracy.
Historically, regression models -- be they linear or complex neural networks -- have mainly used `tabular' data composed of numeric or nominal categorical variables.
However, as companies and institutes increasingly collect more text data, and language models substantially improve every year, there is growing interest in how to leverage text for regression.
This thesis studies using text-only data as well as combining text and tabular data for regression, applied to price prediction.
We begin by discussing an often overlooked aspect of tabular modelling: missing value handling.
A large, real UK insurer dataset is used to study the impact of missing value handling on the performance of gradient boosted trees and deep learning models, two contemporary tabular modelling approaches.
In doing so we establish Catboost as a strong tabular model, relevant for benchmarking joint text-tabular models, but emphasize the importance of missing values.
We then conduct a study of text-only regression.
Text-only regression requires a pipeline of preprocessing, tokenization, featurization and modelling steps.
We rigorously analyze choices for these steps, where the literature lacks systematic work.
Among other results, for the first time, our study quantifies the relative importance of the model choice compared to choice of steps before modelling.
We find that choice of the upstream steps contributes variance in performance comparable to model choice.
We conclude with a study of joint text-tabular modelling: studying an ensemble of a tabular-only Catboost and a text-only neural network.
This ensemble is compared to existing methods.
Overall, we conclude that although this ensemble does not consistently have the best accuracy, it is simple, robust to data with low signal in the text modality and achieves good accuracy.
Version
Open Access
Date Issued
2025-03-14
Date Awarded
01/08/2025
License URL
Advisor
Adams, Niall
Webster, Kevin
Sponsor
Dunnhumby Limited (Firm)
Grant Number
EP/S023151/1
Publisher Department
Department of Mathematics
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
