Correlation set analysis: detecting active regulators in disease populations using prior causal knowledge
File(s)BMCBioinformatics-Correlations(2012).pdf (1.14 MB)
Published version
Author(s)
Type
Journal Article
Abstract
Background: Identification of active causal regulators is a crucial problem in understanding mechanism of diseases
or finding drug targets. Methods that infer causal regulators directly from primary data have been proposed and
successfully validated in some cases. These methods necessarily require very large sample sizes or a mix of
different data types. Recent studies have shown that prior biological knowledge can successfully boost a method’s
ability to find regulators.
Results: We present a simple data-driven method, Correlation Set Analysis (CSA), for comprehensively detecting
active regulators in disease populations by integrating co-expression analysis and a specific type of literaturederived causal relationships. Instead of investigating the co-expression level between regulators and their
regulatees, we focus on coherence of regulatees of a regulator. Using simulated datasets we show that our
method performs very well at recovering even weak regulatory relationships with a low false discovery rate. Using
three separate real biological datasets we were able to recover well known and as yet undescribed, active
regulators for each disease population. The results are represented as a rank-ordered list of regulators, and reveals
both single and higher-order regulatory relationships.
Conclusions: CSA is an intuitive data-driven way of selecting directed perturbation experiments that are relevant
to a disease population of interest and represent a starting point for further investigation. Our findings
demonstrate that combining co-expression analysis on regulatee sets with a literature-derived network can
successfully identify causal regulators and help develop possible hypothesis to explain disease progression.
or finding drug targets. Methods that infer causal regulators directly from primary data have been proposed and
successfully validated in some cases. These methods necessarily require very large sample sizes or a mix of
different data types. Recent studies have shown that prior biological knowledge can successfully boost a method’s
ability to find regulators.
Results: We present a simple data-driven method, Correlation Set Analysis (CSA), for comprehensively detecting
active regulators in disease populations by integrating co-expression analysis and a specific type of literaturederived causal relationships. Instead of investigating the co-expression level between regulators and their
regulatees, we focus on coherence of regulatees of a regulator. Using simulated datasets we show that our
method performs very well at recovering even weak regulatory relationships with a low false discovery rate. Using
three separate real biological datasets we were able to recover well known and as yet undescribed, active
regulators for each disease population. The results are represented as a rank-ordered list of regulators, and reveals
both single and higher-order regulatory relationships.
Conclusions: CSA is an intuitive data-driven way of selecting directed perturbation experiments that are relevant
to a disease population of interest and represent a starting point for further investigation. Our findings
demonstrate that combining co-expression analysis on regulatee sets with a literature-derived network can
successfully identify causal regulators and help develop possible hypothesis to explain disease progression.
Date Issued
2012-03-23
Date Acceptance
2012-03-23
Citation
BMC Bioinformatics, 2012, 13 (1), pp.46-46
ISSN
1471-2105
Publisher
Springer Science and Business Media LLC
Start Page
46
End Page
46
Journal / Book Title
BMC Bioinformatics
Volume
13
Issue
1
Copyright Statement
© 2012 Huang et al; licensee BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons
Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in
any medium, provided the original work is properly cited.
Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in
any medium, provided the original work is properly cited.
License URL
Identifier
https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-13-46
Subjects
Bioinformatics
01 Mathematical Sciences
06 Biological Sciences
08 Information and Computing Sciences
Publication Status
Published
Date Publish Online
2012-03-23