Predicting cancer patient response to chemotherapy using machine learning from small data
File(s)
Author(s)
Du, Hanqin
Type
Thesis
Abstract
Accurately predicting chemotherapy response remains a major challenge in oncology due to tumour heterogeneity and the limited effectiveness of existing biomarkers. Although machine-learning models using tumour omics data have shown promise, most existing studies rely on preclinical cell-line datasets, whose biological and distributional differences from patient tumours often lead to poor generalisation in clinical settings. This study systematically investigates how reliable drug response prediction can be achieved under realistic clinical constraints characterised by small, heterogeneous patient dataset.
Using clinical data from The Cancer Genome Atlas (TCGA), we first evaluated traditional machine learning models across multiple omics modalities in temozolomide-treated low-grade glioma setting. Several omics-based models significantly outperformed the established MGMT methylation biomarker, demonstrating that clinically relevant predictive signals can be extracted from limited patient data. Through repeated nested cross-validation and bootstrap bias correction, we reveal that standard cross validation procedures can still overestimate true predictive performance in small clinical datasets.
We then assessed cross-cancer prediction as a strategy for addressing rare cancers. Across intra-cancer, inter-cancer, and mixed-cancer evaluations, most pan-cancer models failed to achieve robust performance in the absence of target-cancer samples. Only a small subset of drug–cancer pairs consistently exceeded random baselines, and predictability was found to depend more on tumour type than sample size or omics similarity.
Finally, we evaluated several transfer-learning strategies, including biologically informed feature extraction, direct transfer of preclinical models to clinical data, and a conservative hybrid approach that integrates predictions from cell-line–trained models into clinical machine-learning models. The hybrid approach consistently improved predictive performance and model stability across drugs.
Overall, this work demonstrates that evaluation design is as critical as model choice in small clinical datasets and that hybrid transfer-learning approaches provide a more reliable path toward clinically meaningful drug response prediction.
Using clinical data from The Cancer Genome Atlas (TCGA), we first evaluated traditional machine learning models across multiple omics modalities in temozolomide-treated low-grade glioma setting. Several omics-based models significantly outperformed the established MGMT methylation biomarker, demonstrating that clinically relevant predictive signals can be extracted from limited patient data. Through repeated nested cross-validation and bootstrap bias correction, we reveal that standard cross validation procedures can still overestimate true predictive performance in small clinical datasets.
We then assessed cross-cancer prediction as a strategy for addressing rare cancers. Across intra-cancer, inter-cancer, and mixed-cancer evaluations, most pan-cancer models failed to achieve robust performance in the absence of target-cancer samples. Only a small subset of drug–cancer pairs consistently exceeded random baselines, and predictability was found to depend more on tumour type than sample size or omics similarity.
Finally, we evaluated several transfer-learning strategies, including biologically informed feature extraction, direct transfer of preclinical models to clinical data, and a conservative hybrid approach that integrates predictions from cell-line–trained models into clinical machine-learning models. The hybrid approach consistently improved predictive performance and model stability across drugs.
Overall, this work demonstrates that evaluation design is as critical as model choice in small clinical datasets and that hybrid transfer-learning approaches provide a more reliable path toward clinically meaningful drug response prediction.
Version
Open Access
Date Issued
2026-01-05
Date Awarded
2026-06-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Ballester, Pedro
Publisher Department
Department of Bioengineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
