Assessing the quality and readability of patient information available for shockwave lithotripsy: Human versus artificial intelligence comparison of resources and quality measurements - A benchmarking study from European association of urology (EAU) endourology.
Ravanth Baskaran, Omar Al-Gholmy, Wissam Kamal, Carlotta Nedbal, Frederic Panthier, Olivier Traxer, Naeem Bhojani, Joe Philip, Bhaskar K Somani
{"title":"Assessing the quality and readability of patient information available for shockwave lithotripsy: Human versus artificial intelligence comparison of resources and quality measurements - A benchmarking study from European association of urology (EAU) endourology.","authors":"Ravanth Baskaran, Omar Al-Gholmy, Wissam Kamal, Carlotta Nedbal, Frederic Panthier, Olivier Traxer, Naeem Bhojani, Joe Philip, Bhaskar K Somani","doi":"10.4103/ua.ua_42_26","DOIUrl":null,"url":null,"abstract":"<p><p>Shockwave lithotripsy (SWL) is a noninvasive method for removing stones, primarily used for small, uncomplicated urinary stones, as the number of kidney stone cases continues to rise. With rising patient numbers, online information must be accurate and easy to understand. We assessed the quality and readability of SWL information available online and compared it with leaflets created by generative artificial intelligence (AI). The top 20 search results for four SWL-related terms on Google, Yahoo, and Bing were collected and duplicates removed. Generative Artificial Intelligence (GAI) platforms such as ChatGPT, Gemini, and DeepSeek generated a patient leaflet. Two authors evaluated quality using DISCERN and JAMA scores and readability using Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL). Later, ChatGPT reviewed the AI-generated information with DISCERN. A total of 68 websites were included. The quality assessments showed average websites with AI-leaflet DISCERN scores of 52.4 (±8.91) and 58.7 (±5.03), JAMA Benchmark scores of 2.22 (±1.10) and 0, FRES of 52.1 (±13.8) and 51.5 (±8.05), and FKGL scores of 9.04 (±2.36) and 8.93 (±1.10). No significant differences were seen in DISCERN (<i>P</i> = 0.224), FRES (<i>P</i> = 0.652), and FKGL (<i>P</i> = 0.866), but differences in JAMA scores were significant (<i>P</i> = 0.0010). A difference was also noted in the author-rated and the AI-rated DISCERN scores. There is a deficiency in high-quality SWL information that meets global readability standards. As more individuals use GAI-resources trained on online data, it is crucial to improve their quality for patients and carers. Doing so enables them to make informed decisions and promotes their well-being during treatment, resulting in improved comfort and outcomes.</p>","PeriodicalId":23633,"journal":{"name":"Urology Annals","volume":"18 3","pages":"215-224"},"PeriodicalIF":0.9000,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13441066/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Urology Annals","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.4103/ua.ua_42_26","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/7/13 0:00:00","PubModel":"Epub","JCR":"Q4","JCRName":"UROLOGY & NEPHROLOGY","Score":null,"Total":0}
引用次数: 0
Abstract
Shockwave lithotripsy (SWL) is a noninvasive method for removing stones, primarily used for small, uncomplicated urinary stones, as the number of kidney stone cases continues to rise. With rising patient numbers, online information must be accurate and easy to understand. We assessed the quality and readability of SWL information available online and compared it with leaflets created by generative artificial intelligence (AI). The top 20 search results for four SWL-related terms on Google, Yahoo, and Bing were collected and duplicates removed. Generative Artificial Intelligence (GAI) platforms such as ChatGPT, Gemini, and DeepSeek generated a patient leaflet. Two authors evaluated quality using DISCERN and JAMA scores and readability using Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL). Later, ChatGPT reviewed the AI-generated information with DISCERN. A total of 68 websites were included. The quality assessments showed average websites with AI-leaflet DISCERN scores of 52.4 (±8.91) and 58.7 (±5.03), JAMA Benchmark scores of 2.22 (±1.10) and 0, FRES of 52.1 (±13.8) and 51.5 (±8.05), and FKGL scores of 9.04 (±2.36) and 8.93 (±1.10). No significant differences were seen in DISCERN (P = 0.224), FRES (P = 0.652), and FKGL (P = 0.866), but differences in JAMA scores were significant (P = 0.0010). A difference was also noted in the author-rated and the AI-rated DISCERN scores. There is a deficiency in high-quality SWL information that meets global readability standards. As more individuals use GAI-resources trained on online data, it is crucial to improve their quality for patients and carers. Doing so enables them to make informed decisions and promotes their well-being during treatment, resulting in improved comfort and outcomes.