Auto-CORPus: automated and consistent outputs from research publications
File(s)Auto_CORPus_bioRxiv.pdf (1.87 MB)
Working paper
OA Location
Author(s)
Hu, Yan
Sun, Shujian
Rowlands, Thomas
Beck, Tim
Posma, Joram Matthias
Type
Working Paper
Abstract
Motivation: The availability of improved natural language processing (NLP) algorithms and models enable researchers to analyse larger corpora using open source tools. Text mining of biomedical literature is one area for which NLP has been used in recent years with large untapped potential. However, in order to generate corpora that can be analyzed using machine learning NLP algorithms, these need to be standardized. Summarizing data from literature to be stored into databases typically requires manual curation, especially for extracting data from result tables. Results: We present here an automated pipeline that cleans HTML files from biomedical literature. The output is a single JSON file that contains the text for each section, table data in machine-readable format and lists of phenotypes and abbreviations found in the article. We analyzed a total of 2,441 Open Access articles from PubMed Central, from both Genome-Wide and Metabolome-Wide Association Studies, and developed a model to standardize the section headers based on the Information Artifact Ontology. Extraction of table data was developed on PubMed articles and fine-tuned using the equivalent publisher versions. Availability: The Auto-CORPus package is freely available with detailed instructions from Github at https://github.com/jmp111/AutoCORPus/.
Date Issued
2021-01-08
Citation
2021
Publisher
bioRxiv
Copyright Statement
© 2021 The Author(s). It is made
available under a CC-BY-NC-ND 4.0 International license.
available under a CC-BY-NC-ND 4.0 International license.
Sponsor
Medical Research Council (MRC)
Medical Research Council
Identifier
https://www.biorxiv.org/content/10.1101/2021.01.08.425887v1
Grant Number
MR/S004033/1
MR/S004033/1
Subjects
information artefact ontology
natural language processing
text standardization
Notes
This is the pre-print version, article has been submitted to Bioinformatics on 9 Jan 2021.
Publication Status
Published