Rank-based hypothesis testing in the presence of missing data
File(s)
Author(s)
Zeng, Yijin
Type
Thesis
Abstract
Nearly all hypothesis testing methods are designed solely for data that are completely observed. Despite the prevalence of missing data problems, no standard approach exists to address this issue. This thesis proposes several rank-based hypothesis testing methods for distinct, univariate data in the presence of missing data, without making any assumptions of the missing data.
The importance and difficulty of addressing the missing data problem properly motivate the research in this thesis. Standard hypothesis testing methods cannot be implemented directly in the presence of missing data. Common missing data methods, such as case deletion and imputation methods, risk increasing the Type I error beyond a specific significance level. Meanwhile, the vast literature on this topic addresses this issue only under restrictive assumptions by constraining the missing data to specific missingness mechanisms.
The key contributions of this thesis are the development of several rank-based hypothesis testing methods, which, unlike existing methods, can accommodate the missing values in a manner without making any assumptions of the missing data. We prove that our proposed methods control the Type I error regardless of the values of the missing data. Simulation results demonstrate that these methods have good statistical power, typically when less than 15–20% of the data are missing, while all the other methods fail to control the Type I error.
The importance and difficulty of addressing the missing data problem properly motivate the research in this thesis. Standard hypothesis testing methods cannot be implemented directly in the presence of missing data. Common missing data methods, such as case deletion and imputation methods, risk increasing the Type I error beyond a specific significance level. Meanwhile, the vast literature on this topic addresses this issue only under restrictive assumptions by constraining the missing data to specific missingness mechanisms.
The key contributions of this thesis are the development of several rank-based hypothesis testing methods, which, unlike existing methods, can accommodate the missing values in a manner without making any assumptions of the missing data. We prove that our proposed methods control the Type I error regardless of the values of the missing data. Simulation results demonstrate that these methods have good statistical power, typically when less than 15–20% of the data are missing, while all the other methods fail to control the Type I error.
Version
Open Access
Date Issued
2025-09-26
Date Awarded
2026-03-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Bodenham, Dean
Adams, Niall
Sponsor
Imperial College London
Engineering and Physical Sciences Research Council
Publisher Department
Department of Mathematics
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
