Relation Extraction for diet, non-Communicable Disease and biomarker associations (RECoDe): a CoDiet study
File(s) ECCB2026_RECoDe.pdf (560.14 KB)
Accepted version
Author(s)
Type
Journal Article
Abstract
Diet plays a critical role in human health, with growing evidence linking dietary habits to disease outcomes. However, extracting structured
dietary knowledge from biomedical literature remains challenging due to the lack of dedicated relation extraction datasets. To address this
gap, we introduce RECoDe, a novel relation extraction (RE) dataset designed specifically for diet, disease, and related biomedical entities.
RECoDe captures a diverse set of relation types, including a broad spectrum of positive association patterns and explicit negative examples,
with over 5,000 human-annotated instances validated by up to five independent annotators. Furthermore, we benchmark various natural
language processing (NLP) RE models, including BERT-based architectures and enhanced prompting techniques with locally deployed large
language models (LLMs) to improve classification performance on underrepresented relation types. The best performing model was gpt-oss-
20B, a locally-deployed open-weight LLM, achieving an F1-score of 64% (macro) for multi-class classification and 92% for binary classification
using a hierarchical prompting strategy with a separate reflection step built in. To demonstrate the practical utility of RECoDe, we introduce
the Contextual Co-occurrence Summarisation (CoCoS) framework, which aggregates sentence-level relation extractions into document-level
summaries and further integrates evidence across multiple documents. CoCoS produces effect estimates consistent with established dietary
knowledge, demonstrating its validity as a general framework for systematic evidence synthesis.
Availability: The code, models, and dataset are publicly available at https://github.com/omicsNLP/RECoDe.
dietary knowledge from biomedical literature remains challenging due to the lack of dedicated relation extraction datasets. To address this
gap, we introduce RECoDe, a novel relation extraction (RE) dataset designed specifically for diet, disease, and related biomedical entities.
RECoDe captures a diverse set of relation types, including a broad spectrum of positive association patterns and explicit negative examples,
with over 5,000 human-annotated instances validated by up to five independent annotators. Furthermore, we benchmark various natural
language processing (NLP) RE models, including BERT-based architectures and enhanced prompting techniques with locally deployed large
language models (LLMs) to improve classification performance on underrepresented relation types. The best performing model was gpt-oss-
20B, a locally-deployed open-weight LLM, achieving an F1-score of 64% (macro) for multi-class classification and 92% for binary classification
using a hierarchical prompting strategy with a separate reflection step built in. To demonstrate the practical utility of RECoDe, we introduce
the Contextual Co-occurrence Summarisation (CoCoS) framework, which aggregates sentence-level relation extractions into document-level
summaries and further integrates evidence across multiple documents. CoCoS produces effect estimates consistent with established dietary
knowledge, demonstrating its validity as a general framework for systematic evidence synthesis.
Availability: The code, models, and dataset are publicly available at https://github.com/omicsNLP/RECoDe.
Date Acceptance
2026-06-22
Citation
Bioinformatics
ISSN
1367-4803
Publisher
Oxford University Press
Journal / Book Title
Bioinformatics
Copyright Statement
Copyright This paper is embargoed until publication. Once published the author’s accepted manuscript will be made available under a CC-BY License in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy).
License URL
Subjects
Biomedical Text Mining
Diet-Disease Relationships
Knowledge Graph
Large Language Models
Relation Extraction
Publication Status
Accepted
Article Number
btag439
