Defining the Estimated Core Genome of Bacterial Populations Using a Bayesian Decision Model
Author(s)
Type
Journal Article
Abstract
The bacterial core genome is of intense interest and the volume of whole genome sequence data in the public domain
available to investigate it has increased dramatically. The aim of our study was to develop a model to estimate the bacterial
core genome from next-generation whole genome sequencing data and use this model to identify novel genes associated
with important biological functions. Five bacterial datasets were analysed, comprising 2096 genomes in total. We developed
a Bayesian decision model to estimate the number of core genes, calculated pairwise evolutionary distances (p-distances)
based on nucleotide sequence diversity, and plotted the median p-distance for each core gene relative to its genome
location. We designed visually-informative genome diagrams to depict areas of interest in genomes. Case studies
demonstrated how the model could identify areas for further study, e.g. 25% of the core genes with higher sequence
diversity in the
Campylobacter jejuni
and
Neisseria meningitidis
genomes encoded hypothetical proteins. The core gene with
the highest p-distance value in
C. jejuni
was annotated in the reference genome as a putative hydrolase, but further work
revealed that it shared sequence homology with beta-lactamase/metallo-beta-lactamases (enzymes that provide resistance
to a range of broad-spectrum antibiotics) and thioredoxin reductase genes (which reduce oxidative stress and are essential
for DNA replication) in other
C. jejuni
genomes. Our Bayesian model of estimating the core genome is principled, easy to use
and can be applied to large genome datasets. This study also highlighted the lack of knowledge currently available for many
core genes in bacterial genomes of significant global public health importance.
available to investigate it has increased dramatically. The aim of our study was to develop a model to estimate the bacterial
core genome from next-generation whole genome sequencing data and use this model to identify novel genes associated
with important biological functions. Five bacterial datasets were analysed, comprising 2096 genomes in total. We developed
a Bayesian decision model to estimate the number of core genes, calculated pairwise evolutionary distances (p-distances)
based on nucleotide sequence diversity, and plotted the median p-distance for each core gene relative to its genome
location. We designed visually-informative genome diagrams to depict areas of interest in genomes. Case studies
demonstrated how the model could identify areas for further study, e.g. 25% of the core genes with higher sequence
diversity in the
Campylobacter jejuni
and
Neisseria meningitidis
genomes encoded hypothetical proteins. The core gene with
the highest p-distance value in
C. jejuni
was annotated in the reference genome as a putative hydrolase, but further work
revealed that it shared sequence homology with beta-lactamase/metallo-beta-lactamases (enzymes that provide resistance
to a range of broad-spectrum antibiotics) and thioredoxin reductase genes (which reduce oxidative stress and are essential
for DNA replication) in other
C. jejuni
genomes. Our Bayesian model of estimating the core genome is principled, easy to use
and can be applied to large genome datasets. This study also highlighted the lack of knowledge currently available for many
core genes in bacterial genomes of significant global public health importance.
Date Issued
2014-08-21
Date Acceptance
2014-07-01
Citation
PLOS COMPUTATIONAL BIOLOGY, 2014, 10 (8)
ISSN
1553-734X
Publisher
PUBLIC LIBRARY OF SCIENCE
Journal / Book Title
PLOS COMPUTATIONAL BIOLOGY
Volume
10
Issue
8
Copyright Statement
© 2014 van Tonder et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits
unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000341573600037&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Subjects
Science & Technology
Life Sciences & Biomedicine
Biochemical Research Methods
Mathematical & Computational Biology
Biochemistry & Molecular Biology
STAPHYLOCOCCUS-AUREUS
PHYLOGENETIC NETWORKS
PAN-GENOME
EVOLUTION
STRAIN
EPIDEMIOLOGY
MECHANISM
ELEMENTS
SPREAD
GENES
Bacterial Proteins
Bayes Theorem
Campylobacter jejuni
Databases, Genetic
Genome, Bacterial
Genomics
Models, Genetic
Neisseria meningitidis
06 Biological Sciences
08 Information And Computing Sciences
01 Mathematical Sciences
Bioinformatics
Publication Status
Published
Article Number
ARTN e1003788