A resource-efficient tool for mixed model association analysis of large-scale data
File(s)598110v2.full.pdf (1.44 MB)
Accepted version
Author(s)
Type
Journal Article
Abstract
The genome-wide association study (GWAS) has been widely used as an experimental design to detect associations between genetic variants and a phenotype. Two major confounding factors, population stratification and relatedness, could potentially lead to inflated GWAS test statistics and hence to spurious associations. Mixed linear model (MLM)-based approaches can be used to account for sample structure. However, genome-wide association (GWA) analyses in biobank samples such as the UK Biobank (UKB) often exceed the capability of most existing MLM-based tools especially if the number of traits is large. Here, we develop an MLM-based tool (fastGWA) that controls for population stratification by principal components and for relatedness by a sparse genetic relationship matrix for GWA analyses of biobank-scale data. We demonstrate by extensive simulations that fastGWA is reliable, robust and highly resource-efficient. We then apply fastGWA to 2,173 traits on array-genotyped and imputed samples from 456,422 individuals and to 2,048 traits on whole-exome-sequenced samples from 46,191 individuals in the UKB.
Date Issued
2019-12-01
Date Acceptance
2019-10-16
Citation
Nature Genetics, 2019, 51 (12), pp.1749-1755
ISSN
1061-4036
Publisher
Springer Science and Business Media LLC
Start Page
1749
End Page
1755
Journal / Book Title
Nature Genetics
Volume
51
Issue
12
Copyright Statement
© 2020 Springer Nature Limited
Identifier
https://www.nature.com/articles/s41588-019-0530-8
Subjects
06 Biological Sciences
11 Medical and Health Sciences
Developmental Biology
Publication Status
Published
Date Publish Online
2019-11-25