Predictive accuracy of peptide–target interaction models in drug discovery: a systematic review and meta-analysis
File(s) PTI Binding Accuracy SR and MA.pdf (980.98 KB)
Accepted version
Author(s)
Waldock, William
Type
Journal Article
Abstract
Background: Peptides represent promising therapeutic agents due to their high specificity, biocompatibility, and
capacity to modulate protein–protein interactions. However, the field faces critical challenges: inconsistent
evaluation metrics, heterogeneous datasets, and poor reproducibility, which together undermine objective model comparison and benchmarking.
Objectives: To systematically review and quantitatively assess the predictive performance of machine learning
(ML)-based peptide–target interaction (PTI) models, with emphasis on commonly reported metrics including area
under the curve (AUC), concordance index (CI), and Precision.
Methods: We conducted a systematic literature search across PubMed, arXiv, and Cochrane databases for
studies published through August 1, 2025. Inclusion criteria required original ML-based peptide–target prediction models with quantitative performance metrics.
Results: Twenty-three studies met inclusion criteria. Fourteen reported AUC values (pooled estimate: 0.87, 95% Confidence Interval: 0.83–0.90), five reported Concordance Index values (0.90, 95% Confidence Interval: 0.89–39 0.91), and six reported precision (0.75, 95% Confidence Interval: 0.69–0.81). Top-performing models predominantly employed transformer or graph neural network architectures with structural input features. Critical limitations included inconsistent reporting practices, infrequent external validation, and limited data/code availability. The observation that the twenty-three studies present challenges to a pooled estimate is itself a consequence of the reporting practices we document, and is why we advance STRIDE.
Conclusions: ML models demonstrate strong potential for PTI prediction, with leading approaches achieving
robust classification and ranking performance. Nevertheless, progress is hindered by non-standardised
evaluation metrics, limited transparency, and insufficient reproducibility. We introduce the STRIDE framework
(encompassing Standardization, Transparency, Representativeness, Integration, Discovery, and Evidence) to establish rigorous evaluation standards, enhance methodological reproducibility, and support future evaluation of clinical applicability.
capacity to modulate protein–protein interactions. However, the field faces critical challenges: inconsistent
evaluation metrics, heterogeneous datasets, and poor reproducibility, which together undermine objective model comparison and benchmarking.
Objectives: To systematically review and quantitatively assess the predictive performance of machine learning
(ML)-based peptide–target interaction (PTI) models, with emphasis on commonly reported metrics including area
under the curve (AUC), concordance index (CI), and Precision.
Methods: We conducted a systematic literature search across PubMed, arXiv, and Cochrane databases for
studies published through August 1, 2025. Inclusion criteria required original ML-based peptide–target prediction models with quantitative performance metrics.
Results: Twenty-three studies met inclusion criteria. Fourteen reported AUC values (pooled estimate: 0.87, 95% Confidence Interval: 0.83–0.90), five reported Concordance Index values (0.90, 95% Confidence Interval: 0.89–39 0.91), and six reported precision (0.75, 95% Confidence Interval: 0.69–0.81). Top-performing models predominantly employed transformer or graph neural network architectures with structural input features. Critical limitations included inconsistent reporting practices, infrequent external validation, and limited data/code availability. The observation that the twenty-three studies present challenges to a pooled estimate is itself a consequence of the reporting practices we document, and is why we advance STRIDE.
Conclusions: ML models demonstrate strong potential for PTI prediction, with leading approaches achieving
robust classification and ranking performance. Nevertheless, progress is hindered by non-standardised
evaluation metrics, limited transparency, and insufficient reproducibility. We introduce the STRIDE framework
(encompassing Standardization, Transparency, Representativeness, Integration, Discovery, and Evidence) to establish rigorous evaluation standards, enhance methodological reproducibility, and support future evaluation of clinical applicability.
Date Acceptance
2026-08-25
Citation
Frontiers in Drug Discovery
ISSN
2674-0338
Publisher
Frontiers Media S.A.
Journal / Book Title
Frontiers in Drug Discovery
Copyright Statement
Copyright This paper is embargoed until publication. Once published the Version of Record (VoR) will be available on immediate open access.
License URL
Publication Status
Accepted
