Comparing the readability of AI-generated and society-authored patient information leaflets in orthopaedics
Author(s)
Kharma, Nasir
Gill, Saran Singh
Shanmugaratnam, Chayan
Gupte, Chinmay
Type
Journal Article
Abstract
Introduction
Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics.
Methods
A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade
levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at p < 0.05.
Results
Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories (p < 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2-10.6] vs. 8.7 [7.8-9.6], p < 0.01). FRE was consistently lower in AI-generated texts across all categories (p < 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all p < 0.01), indicating greater reading complexity.
Conclusion
AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.
Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics.
Methods
A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade
levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at p < 0.05.
Results
Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories (p < 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2-10.6] vs. 8.7 [7.8-9.6], p < 0.01). FRE was consistently lower in AI-generated texts across all categories (p < 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all p < 0.01), indicating greater reading complexity.
Conclusion
AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.
Date Acceptance
2026-07-13
Citation
Journal of Orthopaedic Surgery and Research
ISSN
1749-799X
Publisher
BMC
Journal / Book Title
Journal of Orthopaedic Surgery and Research
Copyright Statement
Copyright © 2026 Copyright Owner. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Publication Status
Accepted
