Seeing through experts' eyes: a foundational vision-language model trained on radiologists' gaze and reasoning
File(s) s44387-026-00136-9_reference.pdf (16.04 MB)
Published version (in press)
Author(s)
Type
Journal Article
Abstract
Large-scale vision-language models have shown promise in automating chest X-ray interpretation. However, their clinical utility remains limited, since most systems optimize for semantic information rather than emulating how experts visually examine and interpret medical images. As a result, current models often overlook critical findings, misrepresent anatomical context, or diverge from established diagnostic workflows. Radiologists, by contrast, follow structured protocols that sequentially assess anatomical regions, reducing missed findings and supporting reliable diagnostic reasoning. We therefore introduce Gaze-X, a vision-language model that leverages radiologists’ eye-tracking data as a behavioral prior for expert diagnostic reasoning. By incorporating gaze trajectories and fixation patterns into pretraining, Gaze-X learns to follow the spatial and temporal structure of radiologist attention. Using a curated dataset of over 30,000 key frames from five radiologists interpreting chest X-rays across diverse disease categories, we show that Gaze-X produces more accurate, interpretable, and expert-consistent outputs across a range of clinically relevant tasks. Unlike autonomous reporting systems, Gaze-X produces verifiable evidence artifacts, including inspection trajectories and finding-linked localized regions, enabling transparent and safe human-AI collaboration. This capability provides a practical route toward more trustworthy, explainable, and diagnostically robust AI for radiology and beyond.
Date Issued
2026-07-13
Date Acceptance
2026-06-25
Citation
npj Artificial Intelligence, 2026
ISSN
3005-1460
Publisher
Nature Portfolio
Journal / Book Title
npj Artificial Intelligence
Copyright Statement
© The Author(s) 2026. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
License URL
Identifier
10.1038/s44387-026-00136-9
Subjects
Lee
K.
Jing
P.
Zhang
Z. et al. Seeing through experts' eyes: a foundational visionlanguage model trained on radiologists' gaze and reasoning. npj
Publication Status
Published online
Date Publish Online
2026-07-13
