Investigation into the annotation of protocol sequencing steps in the sequence read archive
Author(s)
Alnasir, Jamie
Shanahan, Hugh P
Type
Journal Article
Abstract
BACKGROUND: The workflow for the production of high-throughput sequencing data from nucleic acid samples is complex. There are a series of protocol steps to be followed in the preparation of samples for next-generation sequencing. The quantification of bias in a number of protocol steps, namely DNA fractionation, blunting, phosphorylation, adapter ligation and library enrichment, remains to be determined. RESULTS: We examined the experimental metadata of the public repository Sequence Read Archive (SRA) in order to ascertain the level of annotation of important sequencing steps in submissions to the database. Using SQL relational database queries (using the SRAdb SQLite database generated by the Bioconductor consortium) to search for keywords commonly occurring in key preparatory protocol steps partitioned over studies, we found that 7.10%, 5.84% and 7.57% of all records (fragmentation, ligation and enrichment, respectively), had at least one keyword corresponding to one of the three protocol steps. Only 4.06% of all records, partitioned over studies, had keywords for all three steps in the protocol (5.58% of all SRA records). CONCLUSIONS: The current level of annotation in the SRA inhibits systematic studies of bias due to these protocol steps. Downstream from this, meta-analyses and comparative studies based on these data will have a source of bias that cannot be quantified at present.
Date Issued
2015-05-09
Date Acceptance
2015-04-28
Citation
GigaScience, 2015, 4 (1)
ISSN
2047-217X
Publisher
BioMed Central
Journal / Book Title
GigaScience
Volume
4
Issue
1
Copyright Statement
© 2015 Alnasir and Shanahan; licensee BioMed Central. This is an Open Access article distributed under the terms of theCreative Commons Attribution License (http://creativecommons.org/licenses/by/4.0), which permits unrestricted use,distribution, and reproduction in any medium, provided the original work is properly credited. The Creative Commons PublicDomain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in thisarticle, unless otherwise stated.
Identifier
https://www.ncbi.nlm.nih.gov/pubmed/25960871
PII: 64
Subjects
Annotation
Enrichment
Experiment
Fragmentation
Ligation
Metadata
Next-generation sequencing
Protocol
Computational Biology
Internet
User-Computer Interface
Publication Status
Published
Coverage Spatial
United States
Article Number
ARTN 23
