Multimodal learning in healthcare
File(s)
Author(s)
Liu, Che
Type
Thesis or dissertation
Abstract
This thesis investigates multimodal learning in healthcare by exploring pairwise combinations of chest X-ray images, chest CT scans, electrocardiograms (ECGs), and clinical reports. Relying on a single modality often limits diagnostic accuracy due to incomplete information and modality-specific biases, and using only one modality can underutilize valuable clinical knowledge present in others. Moreover, supervised learning approaches face significant challenges in medical domains due to the scarcity of annotated data, while self-supervised learning (SSL) methods remain largely constrained to single-modality settings, failing to take advantage of cross-modal relationships.
To address these limitations, we first propose a multimodal representation learning framework that links 2D chest X-rays with paired radiology reports. This enables the model to benefit from clinical language supervision, informing visual feature. To reduce reliance on real data, we also investigate the feasibility of training this framework entirely on synthetic image-text pairs generated by modern generative models, providing insights into scalable multimodal pretraining. Secondly, we extend our study to 3D chest CT scans and associated reports. This setting introduces new challenges due to the volumetric nature of CT data and high computational costs. We develop efficient techniques for learning dense 3D visual representations guided by textual context, enabling effective representation learning in this more complex domain. Finally, we move beyond imaging to investigate multimodal learning with ECG signals and corresponding ECG reports. This scenario allows us to explore how structured clinical text can enhance representation learning from physiological time-series data, enabling improved performance in downstream tasks without annotated data.
Together, these studies highlight the versatility and effectiveness of two-modality learning strategies across diverse clinical data types, offering practical approaches to mitigate annotation scarcity, modality imbalance, and limited generalization in medical AI.
To address these limitations, we first propose a multimodal representation learning framework that links 2D chest X-rays with paired radiology reports. This enables the model to benefit from clinical language supervision, informing visual feature. To reduce reliance on real data, we also investigate the feasibility of training this framework entirely on synthetic image-text pairs generated by modern generative models, providing insights into scalable multimodal pretraining. Secondly, we extend our study to 3D chest CT scans and associated reports. This setting introduces new challenges due to the volumetric nature of CT data and high computational costs. We develop efficient techniques for learning dense 3D visual representations guided by textual context, enabling effective representation learning in this more complex domain. Finally, we move beyond imaging to investigate multimodal learning with ECG signals and corresponding ECG reports. This scenario allows us to explore how structured clinical text can enhance representation learning from physiological time-series data, enabling improved performance in downstream tasks without annotated data.
Together, these studies highlight the versatility and effectiveness of two-modality learning strategies across diverse clinical data types, offering practical approaches to mitigate annotation scarcity, modality imbalance, and limited generalization in medical AI.
Version
Open Access
Date Issued
2025-08-16
Date Awarded
2026-08-01
Copyright Statement
Attribution-NonCommercial-ShareAlike 4.0 International Licence (CC BY NC-SA)
Advisor
Arcucci, Rossella
Bai, Wenjia
Shah, Anand
Publisher Department
Department of Earth Science & Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
