Gene ranking and biomarker discovery under correlation
File(s)0902.0751v3.pdf (1 MB)
Accepted version
Author(s)
Zuber, Verena
Strimmer, Korbinian
Type
Journal Article
Abstract
Motivation: Biomarker discovery and gene ranking is a standard task in genomic high-throughput analysis. Typically, the ordering of markers is based on a stabilized variant of the t-score, such as the moderated t or the SAM statistic. However, these procedures ignore gene–gene correlations, which may have a profound impact on the gene orderings and on the power of the subsequent tests.
Results: We propose a simple procedure that adjusts gene-wise t-statistics to take account of correlations among genes. The resulting correlation-adjusted t-scores (‘cat’ scores) are derived from a predictive perspective, i.e. as a score for variable selection to discriminate group membership in two-class linear discriminant analysis. In the absence of correlation the cat score reduces to the standard t-score. Moreover, using the cat score it is straightforward to evaluate groups of features (i.e. gene sets). For computation of the cat score from small sample data, we propose a shrinkage procedure. In a comparative study comprising six different synthetic and empirical correlation structures, we show that the cat score improves estimation of gene orderings and leads to higher power for fixed true discovery rate, and vice versa. Finally, we also illustrate the cat score by analyzing metabolomic data.
Results: We propose a simple procedure that adjusts gene-wise t-statistics to take account of correlations among genes. The resulting correlation-adjusted t-scores (‘cat’ scores) are derived from a predictive perspective, i.e. as a score for variable selection to discriminate group membership in two-class linear discriminant analysis. In the absence of correlation the cat score reduces to the standard t-score. Moreover, using the cat score it is straightforward to evaluate groups of features (i.e. gene sets). For computation of the cat score from small sample data, we propose a shrinkage procedure. In a comparative study comprising six different synthetic and empirical correlation structures, we show that the cat score improves estimation of gene orderings and leads to higher power for fixed true discovery rate, and vice versa. Finally, we also illustrate the cat score by analyzing metabolomic data.
Date Issued
2009-10-15
Date Acceptance
2009-07-22
Citation
Bioinformatics, 2009, 25 (20), pp.2700-2707
ISSN
1367-4803
Publisher
Oxford University Press (OUP)
Start Page
2700
End Page
2707
Journal / Book Title
Bioinformatics
Volume
25
Issue
20
Copyright Statement
© The Author 2009. Published by Oxford University Press. All rights reserved. For Permissions, please email: journals.permissions@oxfordjournals.org. This is a pre-copy-editing, author-produced version
of an article accepted for publication in Bioinformatics following
peer review. The definitive publisher-authenticated version erena Zuber, Korbinian Strimmer, Gene ranking and biomarker discovery under correlation, Bioinformatics, Volume 25, Issue 20, 15 October 2009, Pages 2700–2707 is available online at: https://doi.org/10.1093/bioinformatics/btp460
of an article accepted for publication in Bioinformatics following
peer review. The definitive publisher-authenticated version erena Zuber, Korbinian Strimmer, Gene ranking and biomarker discovery under correlation, Bioinformatics, Volume 25, Issue 20, 15 October 2009, Pages 2700–2707 is available online at: https://doi.org/10.1093/bioinformatics/btp460
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000270685200011&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Subjects
Science & Technology
Life Sciences & Biomedicine
Technology
Physical Sciences
Biochemical Research Methods
Biotechnology & Applied Microbiology
Computer Science, Interdisciplinary Applications
Mathematical & Computational Biology
Statistics & Probability
Biochemistry & Molecular Biology
Computer Science
Mathematics
DISCRIMINANT-ANALYSIS
VARIABLE SELECTION
SHRINKAGE APPROACH
BREAST-CANCER
EXPRESSION
MICROARRAYS
KNOWLEDGE
PROFILES
BAYES
Publication Status
Published
Date Publish Online
2009-07-30