Publication:
Guideline-based, but not error-free: multilingual risks in AI-powered patient counseling on gallstones

Placeholder

Departments

School / College / Institute

Program

KU-Authors

KU Authors

Co-Authors

Erdem, O.
Canbak, T.
Acar, A.
Ceylan, E. M.
Cakit, H.
Basak, F.

Editor & Affiliation

Compiler & Affiliation

Translator

Other Contributor

Date

Language

eng

Embargo Status

N/A

Journal Title

Journal ISSN

Volume Title

Alternative Title

Abstract

Patients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs—particularly across languages—remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. Methods In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0–2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran’s Q with McNemar post-hoc tests. Error types and response length were also examined. Results In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . Conclusion LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance—especially in low-resource languages—should be discouraged.

Source

Publisher

Elsevier

Subject

Computer science, Health care sciences and services, Medical informatics

Citation

Has Part

Source

International Journal of Medical Informatics

Book Series Title

Edition

DOI

10.1016/j.ijmedinf.2026.106341

item.page.datauri

Link

Rights

N/A

Copyrights Note

Creative Commons license

Except where otherwised noted, this item's license is described as N/A

Endorsement

Review

Supplemented By

Referenced By

Related Goal

0

Views

0

Downloads

View PlumX Details