Publication: Guideline-based, but not error-free: multilingual risks in AI-powered patient counseling on gallstones
| dc.contributor.coauthor | Erdem, O. | |
| dc.contributor.coauthor | Canbak, T. | |
| dc.contributor.coauthor | Acar, A. | |
| dc.contributor.coauthor | Ceylan, E. M. | |
| dc.contributor.coauthor | Cakit, H. | |
| dc.contributor.coauthor | Basak, F. | |
| dc.date.accessioned | 2026-08-14T11:24:09Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Patients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs—particularly across languages—remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. Methods In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0–2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran’s Q with McNemar post-hoc tests. Error types and response length were also examined. Results In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . Conclusion LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance—especially in low-resource languages—should be discouraged. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | WOS | |
| dc.description.indexedby | Scopus | |
| dc.description.indexedby | PubMed | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | N/A | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | 84 | |
| dc.identifier.ScopusQuartile | Q1 | |
| dc.identifier.WoSPercentile | 87,4 | |
| dc.identifier.WoSQuartile | Q1 | |
| dc.identifier.doi | 10.1016/j.ijmedinf.2026.106341 | |
| dc.identifier.eissn | 1872-8243 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.issn | 1386-5056 | |
| dc.identifier.pubmed | 41689953 | |
| dc.identifier.scopus | 2-s2.0-105030076162 | |
| dc.identifier.uri | http://doi.org/10.1016/j.ijmedinf.2026.106341 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34460 | |
| dc.identifier.volume | 212 | |
| dc.identifier.wos | 001693746600001 | |
| dc.keywords | Large language models | |
| dc.keywords | Digital health | |
| dc.keywords | Gallstone disease | |
| dc.keywords | Language bias | |
| dc.keywords | ChatGPT | |
| dc.keywords | Guideline concordance | |
| dc.keywords | Gemini | |
| dc.language | eng | |
| dc.publisher | Elsevier | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | International Journal of Medical Informatics | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Computer science | |
| dc.subject | Health care sciences and services | |
| dc.subject | Medical informatics | |
| dc.title | Guideline-based, but not error-free: multilingual risks in AI-powered patient counseling on gallstones | |
| dc.type | Journal Article | |
| dspace.entity.type | Publication |
