Using likelihood-free inference to compare evolutionary dynamics of the protein networks of H. pylori and P. falciparum
Author(s)
Type
Journal Article
Abstract
Gene duplication with subsequent interaction divergence is one of the primary driving forces in the evolution of genetic systems. Yet little is known about the precise mechanisms and the role of duplication divergence in the evolution of protein networks from the prokaryote and eukaryote domains. We developed a novel, model-based approach for Bayesian inference on biological network data that centres on approximate Bayesian computation, or likelihood-free inference. Instead of computing the intractable likelihood of the protein network topology, our method summarizes key features of the network and, based on these, uses a MCMC algorithm to approximate the posterior distribution of the model parameters. This allowed us to reliably fit a flexible mixture model that captures hallmarks of evolution by gene duplication and subfunctionalization to protein interaction network data of Helicobacter pylori and Plasmodium falciparum. The 80% credible intervals for the duplication–divergence component are [0.64, 0.98] for H. pylori and [0.87, 0.99] for P. falciparum. The remaining parameter estimates are not inconsistent with sequence data. An extensive sensitivity analysis showed that incompleteness of PIN data does not largely affect the analysis of models of protein network evolution, and that the degree sequence alone barely captures the evolutionary footprints of protein networks relative to other statistics. Our likelihood-free inference approach enables a fully Bayesian analysis of a complex and highly stochastic system that is otherwise intractable at present. Modelling the evolutionary history of PIN data, it transpires that only the simultaneous analysis of several global aspects of protein networks enables credible and consistent inference to be made from available datasets. Our results indicate that gene duplication has played a larger part in the network evolution of the eukaryote than in the prokaryote, and suggests that single gene duplications with immediate divergence alone may explain more than 60% of biological network data in both domains.
Date Issued
2007-11-30
Date Acceptance
2007-10-05
Citation
PLOS Computational Biology, 2007, 3 (11)
ISSN
1553-734X
Publisher
Public Library of Science
Journal / Book Title
PLOS Computational Biology
Volume
3
Issue
11
Copyright Statement
© 2007 Ratmann et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
License URL
Subjects
Science & Technology
Life Sciences & Biomedicine
Biochemical Research Methods
Mathematical & Computational Biology
Biochemistry & Molecular Biology
BIOCHEMICAL RESEARCH METHODS
MATHEMATICAL & COMPUTATIONAL BIOLOGY
DUPLICATE GENES
COMPLEX NETWORKS
COMPARATIVE GENOMICS
PURIFYING SELECTION
DEATH EVOLUTION
MONTE-CARLO
FAMILY
DIVERGENCE
PRESERVATION
EXPRESSION
Animals
Bacterial Proteins
Biological Evolution
Evolution, Molecular
Gene Duplication
Genetic Variation
Helicobacter pylori
Likelihood Functions
Models, Genetic
Models, Statistical
Plasmodium falciparum
Protozoan Proteins
Signal Transduction
Bioinformatics
06 Biological Sciences
08 Information And Computing Sciences
01 Mathematical Sciences
Publication Status
Published
Article Number
e230