Flexible fitting of PROTAC concentration-response curves with changepoint Gaussian processes
File(s)1-s2.0-S2472555222067600-main.pdf (2.13 MB)
Published version
Author(s)
Type
Journal Article
Abstract
A proteolysis-targeting chimera (PROTAC) is a new technology that marks proteins for degradation in a highly specific
manner. During screening, PROTAC compounds are tested in concentration–response (CR) assays to determine their
potency, and parameters such as the half-maximal degradation concentration (DC50) are estimated from the fitted CR
curves. These parameters are used to rank compounds, with lower DC50 values indicating greater potency. However,
PROTAC data often exhibit biphasic and polyphasic relationships, making standard sigmoidal CR models inappropriate.
A common solution includes manual omitting of points (the so-called masking step), allowing standard models to be
used on the reduced data sets. Due to its manual and subjective nature, masking becomes a costly and nonreproducible
procedure. We therefore used a Bayesian changepoint Gaussian processes model that can flexibly fit both nonsigmoidal
and sigmoidal CR curves without user input. Parameters such as the DC50, maximum effect Dmax, and point of departure
(PoD) are estimated from the fitted curves. We then rank compounds based on one or more parameters and propagate the
parameter uncertainty into the rankings, enabling us to confidently state if one compound is better than another. Hence,
we used a flexible and automated procedure for PROTAC screening experiments. By minimizing subjective decisions, our
approach reduces time and cost and ensures reproducibility of the compound-ranking procedure. The code and data are
provided on GitHub (https://github.com/elizavetasemenova/gp_concentration_response).
manner. During screening, PROTAC compounds are tested in concentration–response (CR) assays to determine their
potency, and parameters such as the half-maximal degradation concentration (DC50) are estimated from the fitted CR
curves. These parameters are used to rank compounds, with lower DC50 values indicating greater potency. However,
PROTAC data often exhibit biphasic and polyphasic relationships, making standard sigmoidal CR models inappropriate.
A common solution includes manual omitting of points (the so-called masking step), allowing standard models to be
used on the reduced data sets. Due to its manual and subjective nature, masking becomes a costly and nonreproducible
procedure. We therefore used a Bayesian changepoint Gaussian processes model that can flexibly fit both nonsigmoidal
and sigmoidal CR curves without user input. Parameters such as the DC50, maximum effect Dmax, and point of departure
(PoD) are estimated from the fitted curves. We then rank compounds based on one or more parameters and propagate the
parameter uncertainty into the rankings, enabling us to confidently state if one compound is better than another. Hence,
we used a flexible and automated procedure for PROTAC screening experiments. By minimizing subjective decisions, our
approach reduces time and cost and ensures reproducibility of the compound-ranking procedure. The code and data are
provided on GitHub (https://github.com/elizavetasemenova/gp_concentration_response).
Date Issued
2021-10
Date Acceptance
2021-06-04
Citation
SLAS Discovery, 2021, 26 (9), pp.1212-1224
ISSN
2472-5552
Publisher
Elsevier
Start Page
1212
End Page
1224
Journal / Book Title
SLAS Discovery
Volume
26
Issue
9
Copyright Statement
© Society for Laboratory
Identifier
https://www.sciencedirect.com/science/article/pii/S2472555222067600?via%3Dihub
Subjects
Bayesian inference
Biochemical Research Methods
Biochemistry & Molecular Biology
Biotechnology & Applied Microbiology
Chemistry
Chemistry, Analytical
compound ranking
concentration-response curve
Gaussian process
Life Sciences & Biomedicine
MODEL
Physical Sciences
PROTAC
Science & Technology
Publication Status
Published
Date Publish Online
2021-09-20