PiCSAR: Probabilistic confidence selection and ranking for reasoning chains
File(s) 4787_PiCSAR_Probabilistic_Conf.pdf (4.06 MB)
Published version
Author(s)
Type
Conference Paper
Abstract
Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate solutions and selecting the one with the highest reward. The key challenge for reasoning tasks is designing a scoring function that can identify correct reasoning chains without access to ground-truth answers. We propose Probabilistic Confidence Selection and Ranking for Reasoning Chains (PiCSAR): a simple, training-free method that scores each candidate generation using the joint log-likelihood of the reasoning and final answer. This method utilises both the scores of the reasoning path (*reasoning confidence*) and the final answer (*answer confidence*). PiCSAR achieves substantial gains across several benchmarks (+11.7 on AIME2024, +9.81 on AIME2025), outperforming baselines with at least 2x fewer samples in 20 out of 25 comparisons. Our analysis reveals that correct reasoning chains exhibit higher reasoning and answer confidence, justifying the effectiveness of PiCSAR.
Date Issued
2026-07-02
Date Acceptance
2026-04-07
Citation
ACL Anthology, 2026, pp.31511-31544
Publisher
Association for Computational Linguistics
Start Page
31511
End Page
31544
Journal / Book Title
ACL Anthology
Copyright Statement
©2026 Association for Computational Linguistics. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.
License URL
Source
Findings of Association for Computational Linguistics (ACL 2026)
Publication Status
Published
Start Date
2026-07-02
Finish Date
2026-07-07
Coverage Spatial
San Diego, California, United States
