Post-reduction inference for confidence sets of models
File(s) asag045.pdf (869.76 KB)
Accepted version
Author(s)
Battey, Heather
Rasines, Daniel G
Tang, Yanbo
Type
Journal Article
Abstract
Sparsity in a regression context makes the model itself an object of interest, pointing to a confidence set of models as the appropriate presentation of evidence. A difficulty in areas such as genomics, where the number of candidate variables is vast, arises from the need for preliminary reduction prior to the assessment of models. The present paper considers a resolution using
inferential separations fundamental to the Fisherian approach to conditional inference, namely, the sufficiency/co-sufficiency separation, and the ancillary/co-ancillary separation. Tests of model adequacy based on such separations do not involve specifying a direction for departure from any postulated model, avoiding issues of calibration that would arise in directed tests from using the same data for reduction and for model assessment. In idealised cases with no nuisance parameters, the separations extract the relevant information without loss or redundancy. The extent to which estimation of nuisance parameters affects this idealisation is illustrated in detail for the normal theory linear regression model, extending immediately to a log-normal accelerated-life model for time-to-event outcomes. As part of the analysis, we introduce a modified version of the refitted cross-validation estimator of Fan et al. (2012), whose distribution theory is tractable in the appropriate conditional sense. The paper concludes, among other things, that in settings where reduction on a reduced sample is unproblematic, sample splitting has high efficiency relative to our conditional analysis, echoing Cox (1975); otherwise it gives miscalibrated confidence sets.
inferential separations fundamental to the Fisherian approach to conditional inference, namely, the sufficiency/co-sufficiency separation, and the ancillary/co-ancillary separation. Tests of model adequacy based on such separations do not involve specifying a direction for departure from any postulated model, avoiding issues of calibration that would arise in directed tests from using the same data for reduction and for model assessment. In idealised cases with no nuisance parameters, the separations extract the relevant information without loss or redundancy. The extent to which estimation of nuisance parameters affects this idealisation is illustrated in detail for the normal theory linear regression model, extending immediately to a log-normal accelerated-life model for time-to-event outcomes. As part of the analysis, we introduce a modified version of the refitted cross-validation estimator of Fan et al. (2012), whose distribution theory is tractable in the appropriate conditional sense. The paper concludes, among other things, that in settings where reduction on a reduced sample is unproblematic, sample splitting has high efficiency relative to our conditional analysis, echoing Cox (1975); otherwise it gives miscalibrated confidence sets.
Date Issued
2026-07-07
Date Acceptance
2026-06-25
Citation
Biometrika, 2026
ISSN
0006-3444
Publisher
Oxford University Press
Journal / Book Title
Biometrika
Copyright Statement
© The Author(s) 2026. Published by Oxford University Press on behalf of Biometrika Trust. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
License URL
Publication Status
Published online
Article Number
asag045
Date Publish Online
2026-07-07
