Publication:
Guideline-based, but not error-free: multilingual risks in AI-powered patient counseling on gallstones

dc.contributor.coauthorErdem, O.
dc.contributor.coauthorCanbak, T.
dc.contributor.coauthorAcar, A.
dc.contributor.coauthorCeylan, E. M.
dc.contributor.coauthorCakit, H.
dc.contributor.coauthorBasak, F.
dc.date.accessioned2026-08-14T11:24:09Z
dc.date.issued2026
dc.description.abstractPatients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs—particularly across languages—remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. Methods In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0–2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran’s Q with McNemar post-hoc tests. Error types and response length were also examined. Results In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . Conclusion LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance—especially in low-resource languages—should be discouraged.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.indexedbyPubMed
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile84
dc.identifier.ScopusQuartileQ1
dc.identifier.WoSPercentile87,4
dc.identifier.WoSQuartileQ1
dc.identifier.doi10.1016/j.ijmedinf.2026.106341
dc.identifier.eissn1872-8243
dc.identifier.embargoN/A
dc.identifier.issn1386-5056
dc.identifier.pubmed41689953
dc.identifier.scopus2-s2.0-105030076162
dc.identifier.urihttp://doi.org/10.1016/j.ijmedinf.2026.106341
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34460
dc.identifier.volume212
dc.identifier.wos001693746600001
dc.keywordsLarge language models
dc.keywordsDigital health
dc.keywordsGallstone disease
dc.keywordsLanguage bias
dc.keywordsChatGPT
dc.keywordsGuideline concordance
dc.keywordsGemini
dc.languageeng
dc.publisherElsevier
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofInternational Journal of Medical Informatics
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectComputer science
dc.subjectHealth care sciences and services
dc.subjectMedical informatics
dc.titleGuideline-based, but not error-free: multilingual risks in AI-powered patient counseling on gallstones
dc.typeJournal Article
dspace.entity.typePublication

Files