Comprehensiveness metrics for automatic evaluation of factual recall in text generation
File(s) 2026.findings-acl.1744.pdf (902.71 KB)
Published version
Author(s)
Dejl, Adam
Barry, James
Pascale, Alessandra
Carnerero-Cano, Javier
Type
Conference Paper
Abstract
Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce outputs that are incomplete or selectively omit key information. In sensitive domains, such omissions can result in significant harm comparable to that posed by factual inaccuracies, including hallucinations. In this study, we address the challenge of evaluating the comprehensiveness of LLM-generated texts, focusing on the detection of missing information or underrepresented viewpoints. We investigate three automated evaluation metrics: (1) an NLI-based method that decomposes texts into atomic statements and uses natural language inference (NLI) to identify missing facts, (2) a Q A-based metric that extracts question-answer pairs and compares responses across sources, and (3) an end-to-end approach that directly identifies missing content using LLMs. Our experiments demonstrate the surprising effectiveness of the simple end-to-end metric compared to more complex metrics, though at the cost of reduced robustness, interpretability and result granularity. We further assess the comprehensiveness of responses from several popular open-weight LLMs when answering user queries based on multiple sources.
Date Issued
2026-07-02
Date Acceptance
2026-07-01
Citation
Findings of the Association for Computational Linguistics: ACL 2026, 2026, pp.34931-34966
Publisher
Association for Computational Linguistics
Start Page
34931
End Page
34966
Journal / Book Title
Findings of the Association for Computational Linguistics: ACL 2026
Copyright Statement
©2026 Association for Computational Linguistics. ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.
License URL
Source
Findings of the Association for Computational Linguistics: ACL 2026
Publication Status
Published
Start Date
2026-07-02
Finish Date
2026-07-07
Coverage Spatial
San Diego, California, United States
