Deep learning for medical ultrasound video understanding and generation
File(s)
Author(s)
Reynaud, Hadrien
Type
Thesis
Abstract
Medical ultrasound imaging, particularly echocardiography, is widely used in healthcare due to its non-invasive nature and real-time capabilities. However, the application of deep learning to echocardiography faces significant challenges, including limited data availability, strict privacy regulations restricting data sharing, and inefficient clinical workflows reliant on manual interpretation. This thesis addresses these challenges by offering three major contributions that bridge deep learning and medical ultrasound imaging.
Firstly, we develop novel transformer-based architectures specifically designed for echocardiogram video analysis, enabling automated detection of end-systolic and end-diastolic frames and accurate estimation of left ventricular ejection fraction. These architectures effectively capture the temporal dynamics of cardiac motion, representing one of the first vision transformer approaches tailored for echocardiography.
Secondly, we present an evolution of generative approaches for echocardiogram synthesis, progressing from GAN-based frameworks to diffusion models and finally to cascaded video diffusion architectures. This progression yields increasingly realistic synthetic echocardiograms with precise control over ejection fraction, culminating in videos indistinguishable from real data by clinical experts.
Finally, we establish frameworks for creating privacy-preserving synthetic echocardiogram datasets that maintain both visual fidelity and real-world utility for downstream tasks. Our latent flow matching approach achieves performance parity between synthetic and real datasets when training downstream models, marking a significant breakthrough in synthetic medical image and video generation.
The methods developed throughout this thesis have broad implications for improving clinical workflows, enhancing research collaboration through privacy-preserving data sharing, and democratizing access to high-quality training data across institutions.
Firstly, we develop novel transformer-based architectures specifically designed for echocardiogram video analysis, enabling automated detection of end-systolic and end-diastolic frames and accurate estimation of left ventricular ejection fraction. These architectures effectively capture the temporal dynamics of cardiac motion, representing one of the first vision transformer approaches tailored for echocardiography.
Secondly, we present an evolution of generative approaches for echocardiogram synthesis, progressing from GAN-based frameworks to diffusion models and finally to cascaded video diffusion architectures. This progression yields increasingly realistic synthetic echocardiograms with precise control over ejection fraction, culminating in videos indistinguishable from real data by clinical experts.
Finally, we establish frameworks for creating privacy-preserving synthetic echocardiogram datasets that maintain both visual fidelity and real-world utility for downstream tasks. Our latent flow matching approach achieves performance parity between synthetic and real datasets when training downstream models, marking a significant breakthrough in synthetic medical image and video generation.
The methods developed throughout this thesis have broad implications for improving clinical workflows, enhancing research collaboration through privacy-preserving data sharing, and democratizing access to high-quality training data across institutions.
Version
Open Access
Date Issued
2025-03-31
Date Awarded
01/09/2025
License URL
Advisor
Kainz, Bernhard
Sponsor
Engineering and Physical Sciences Research Council
Ultromics Ltd
Grant Number
EP/S023283/1
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
