Intelligent algorithms for DNA detection, quantification and multiplexing
File(s)
Author(s)
Moniri, Ahmad
Type
Thesis
Abstract
Nucleic acid detection, referred herein as `DNA detection', is defined as the process of identifying the presence of a specific sequence of bases, and has a vital, far-reaching impact on healthcare applications such as pathogen detection and drug resistance screening. In particular, DNA amplification methods, such as those used in popular COVID-19 tests, present a rapid and affordable means of DNA detection. When performing DNA amplification, there are 3 main goals which are desired: detection, quantification, and multiplexing. In an ideal world, we would like: (i) detection to be rapid and robust; (ii) quantification to be accurate and precise; and (iii) the ability to multiplex several targets in the same chemical reaction, without increasing the complexity of the chemistry or instrument. The aim of this thesis is to develop intelligent algorithms which can be used to improve DNA detection, quantification and multiplexing.
In an effort to achieve the aforementioned goals, several detection platforms have been developed which can be broadly classed as: quantitative, digital, and electronic instruments, corresponding to the 3 Parts in this thesis. This work exhibits a common philosophy throughout: increasing the value of data through algorithms that operate in higher dimensions, compared with existing methods which significantly reduce the dimensionality of data.
In Part I, the use of multiple features from quantitative instruments are investigated by analysing each signal in multidimensions, compared to the traditional unidimensional perspective, where each signal is reduced to a single value. To this end, a new framework is proposed that formalises the process of quantification and leverages the benefits of multiple features using a novel concept called the multidimensional standard curve (MSC). This new method is shown to guarantee enhanced quantification and provides a natural method of outlier detection, solely through analysing existing data from a new perspective - incurring little to no cost. Subsequently, this new framework is extended to perform multiplexing, demonstrating the first data-driven method for simultaneous quantification and multiplexing of nucleic acids in a single fluorescent channel. In particular, this method showed an accuracy of 100.0% (using 11 clinical isolates) when multiplexing 4 of the most prominent carbapenem-resistant genes, which account for 97.1% of confirmed cases in the UK.
In Part II, the use of machine learning methods for multiplexing is explored by considering the vast volume of data from digital instruments. To this end, a supervised machine learning classifier is trained to automatically perform a new form of data-driven multiplexing, referred to as Amplification Curve Analysis (ACA). This demonstrates the first application of machine learning for multiplexing in digital instruments based on amplification curves, leading to breaking the barrier of multiplexing one target per fluorescent channel and final intensity value. Moreover, a mathematical formula based on Multivariate Poisson statistics is derived to theoretically demonstrate the trade-off between multiplexing and quantification, and support evaluating quantification precision in practical settings. Following this work, a novel 3-step machine learning system called Amplification and Melting Curve Analysis (AMCA) is proposed, which combines the kinetic information from amplification curves with thermodynamic information from melting curves - leading to the largest digital multiplex (using cheap intercalating dyes) that has been reported in the literature. In particular, over 100,000 amplification events were analysed, where this method showed an accuracy of 99.3% (for positive events) when multiplexing 9 variants of mobilised colistin resistance, representing a 10.0% increase over existing methods.
In Part III, the use of intelligent algorithms for increasing the robustness of electronic devices based on sensor arrays in highly noisy environments is explored. In particular, recent devices based on Ion-Sensitive Field-Effect Transistor (ISFET) sensors present a new form of data which is correlated in both space and time (similar to a video), and can produce a vast volume of rich data. To this end, a new algorithm is proposed based on modelling sensor drift using adaptive signal processing and clustering sensors based on their behaviour using unsupervised machine learning. This robust algorithm is shown to enable a high precision application such as distinguishing bacterial from viral infections using host signature-based classification. In particular, this new method results in robust values from 100.0% of experiments (N=22), compared with 63.6% from an existing method.
In summary, as inspired by the field of Signal Processing and Machine Learning, this thesis develops new intelligent algorithms for a wide range of platforms, by analysing data in higher dimensions. Through exploring the data from instruments that represent past, present and future diagnostic systems, this thesis lays the foundation for maximising the value of amplification data, which will ultimately lead to better informed decisions across all healthcare applications.
In an effort to achieve the aforementioned goals, several detection platforms have been developed which can be broadly classed as: quantitative, digital, and electronic instruments, corresponding to the 3 Parts in this thesis. This work exhibits a common philosophy throughout: increasing the value of data through algorithms that operate in higher dimensions, compared with existing methods which significantly reduce the dimensionality of data.
In Part I, the use of multiple features from quantitative instruments are investigated by analysing each signal in multidimensions, compared to the traditional unidimensional perspective, where each signal is reduced to a single value. To this end, a new framework is proposed that formalises the process of quantification and leverages the benefits of multiple features using a novel concept called the multidimensional standard curve (MSC). This new method is shown to guarantee enhanced quantification and provides a natural method of outlier detection, solely through analysing existing data from a new perspective - incurring little to no cost. Subsequently, this new framework is extended to perform multiplexing, demonstrating the first data-driven method for simultaneous quantification and multiplexing of nucleic acids in a single fluorescent channel. In particular, this method showed an accuracy of 100.0% (using 11 clinical isolates) when multiplexing 4 of the most prominent carbapenem-resistant genes, which account for 97.1% of confirmed cases in the UK.
In Part II, the use of machine learning methods for multiplexing is explored by considering the vast volume of data from digital instruments. To this end, a supervised machine learning classifier is trained to automatically perform a new form of data-driven multiplexing, referred to as Amplification Curve Analysis (ACA). This demonstrates the first application of machine learning for multiplexing in digital instruments based on amplification curves, leading to breaking the barrier of multiplexing one target per fluorescent channel and final intensity value. Moreover, a mathematical formula based on Multivariate Poisson statistics is derived to theoretically demonstrate the trade-off between multiplexing and quantification, and support evaluating quantification precision in practical settings. Following this work, a novel 3-step machine learning system called Amplification and Melting Curve Analysis (AMCA) is proposed, which combines the kinetic information from amplification curves with thermodynamic information from melting curves - leading to the largest digital multiplex (using cheap intercalating dyes) that has been reported in the literature. In particular, over 100,000 amplification events were analysed, where this method showed an accuracy of 99.3% (for positive events) when multiplexing 9 variants of mobilised colistin resistance, representing a 10.0% increase over existing methods.
In Part III, the use of intelligent algorithms for increasing the robustness of electronic devices based on sensor arrays in highly noisy environments is explored. In particular, recent devices based on Ion-Sensitive Field-Effect Transistor (ISFET) sensors present a new form of data which is correlated in both space and time (similar to a video), and can produce a vast volume of rich data. To this end, a new algorithm is proposed based on modelling sensor drift using adaptive signal processing and clustering sensors based on their behaviour using unsupervised machine learning. This robust algorithm is shown to enable a high precision application such as distinguishing bacterial from viral infections using host signature-based classification. In particular, this new method results in robust values from 100.0% of experiments (N=22), compared with 63.6% from an existing method.
In summary, as inspired by the field of Signal Processing and Machine Learning, this thesis develops new intelligent algorithms for a wide range of platforms, by analysing data in higher dimensions. Through exploring the data from instruments that represent past, present and future diagnostic systems, this thesis lays the foundation for maximising the value of amplification data, which will ultimately lead to better informed decisions across all healthcare applications.
Version
Open Access
Date Issued
2021-01
Date Awarded
2021-05
Copyright Statement
Creative Commons Attribution NonCommercial Licence
Advisor
Georgiou, Pantelakis
Rodriguez-Manzano, Jesus
Sponsor
Engineering and Physical Sciences Research Council
Grant Number
EP/N509486/1
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)