What is a Good Calibration Question?
File(s) Supplementary material.pdf (3.67 MB) What's in a question_Final.docx (12.91 MB)
Supporting information
Accepted version
Author(s)
Hemming, Victoria
Hanea, Anca
Burgman, Mark
Type
Journal Article
Abstract
Weighted aggregation of expert judgements based on their performance on calibration questions may improve mathematically aggregated judgements relative to equal weights. However, obtaining validated, relevant calibration questions can be difficult. If so, should analysts settle for equal weights? Or should they use calibration questions that are easier to obtain but less relevant?
In this paper, we examine what happens to the out-of-sample performance of weighted aggregations of the Classical Model compared to equal weighted aggregations when the set of calibration questions includes many so-called ‘irrelevant’ questions, those that might ordinarily be considered to be outside the domain of the questions of interest.
We find that performance weighted aggregations outperform equal weights on the combined Classical Model (CM) Score, but not on Statistical Accuracy (i.e., calibration). Importantly, there was no appreciable difference in performance when weights were developed on relevant versus irrelevant questions. Experts were unable to adapt their knowledge across vastly different domains, and in-sample validation did not accurately predict out-of-sample performance on irrelevant questions.
We suggest that if relevant calibration questions cannot be found, then analysts should use equal weights, and draw on alternative techniques to improve judgements. Our study also indicates limits to the predictive accuracy of performance weighted aggregation, and the degree to which expertise can be adapted across domains. We note limitations in our study and urge further research into the effect of question type on the reliability of performance weighted aggregations.
In this paper, we examine what happens to the out-of-sample performance of weighted aggregations of the Classical Model compared to equal weighted aggregations when the set of calibration questions includes many so-called ‘irrelevant’ questions, those that might ordinarily be considered to be outside the domain of the questions of interest.
We find that performance weighted aggregations outperform equal weights on the combined Classical Model (CM) Score, but not on Statistical Accuracy (i.e., calibration). Importantly, there was no appreciable difference in performance when weights were developed on relevant versus irrelevant questions. Experts were unable to adapt their knowledge across vastly different domains, and in-sample validation did not accurately predict out-of-sample performance on irrelevant questions.
We suggest that if relevant calibration questions cannot be found, then analysts should use equal weights, and draw on alternative techniques to improve judgements. Our study also indicates limits to the predictive accuracy of performance weighted aggregation, and the degree to which expertise can be adapted across domains. We note limitations in our study and urge further research into the effect of question type on the reliability of performance weighted aggregations.
Date Issued
2022-02-01
Date Acceptance
2021-02-25
Citation
Risk Analysis: an international journal, 2022, 42 (2), pp.264-278
ISSN
0272-4332
Publisher
Wiley
Start Page
264
End Page
278
Journal / Book Title
Risk Analysis: an international journal
Volume
42
Issue
2
Copyright Statement
© 2021 Society for Risk Analysis. This is the peer reviewed version of the following article, which has been published in final form at https://onlinelibrary.wiley.com/doi/10.1111/risa.13725. This article may be used for non-commercial purposes in accordance with Wiley Terms and Conditions for Use of Self-Archived Versions.
Subjects
Science & Technology
Social Sciences
Life Sciences & Biomedicine
Physical Sciences
Public, Environmental & Occupational Health
Mathematics, Interdisciplinary Applications
Social Sciences, Mathematical Methods
Mathematics
Mathematical Methods In Social Sciences
Aggregation
calibration
equal weights
expert judgment
performance weights
EXPERT JUDGMENT
ELICITATION
GUIDE
Aggregation
calibration
equal weights
expert judgment
performance weights
Strategic, Defence & Security Studies
Publication Status
Published
Date Publish Online
2021-04-16
