Publication:
Assessing chatgpt-4o's potential as a support tool in neuroradiology cases: a comparison of performance by using radiologic images

dc.contributor.coauthorSolim, L. A.
dc.contributor.coauthorAtasoy, D.
dc.contributor.coauthorKoch, V.
dc.contributor.coauthorBachir, A. A.
dc.contributor.coauthorYel, I.
dc.contributor.coauthorVogl, T. J.
dc.contributor.departmentKUH (Koç University Hospital)
dc.contributor.kuauthorYüzkan, Sabahattin
dc.contributor.schoolcollegeinstituteKUH (KOÇ UNIVERSITY HOSPITAL)
dc.date.accessioned2026-09-15T10:55:17Z
dc.date.issued2026
dc.description.abstractBackground In clinical settings, particularly in radiology, artificial intelligence and large language models have already demonstrated promising results. They have the potential to become an integral component of radiologists' workflows in the future, and neuroradiology is likely to benefit from their increasing integration into clinical practice. This study aims to evaluate Chat Generative Pre‐trained Transformer 4 Omni (ChatGPT‐4o) as a supportive tool in neuroradiology by assessing its ability to generate accurate radiology reports from imaging and patient history and its impact on radiologists' diagnostic accuracy. Methods Retrospective analysis was conducted using only radiological images and brief patient admission history from 30 neuroradiology cases published publicly between 2015 and 2024. ChatGPT‐4o generated radiology reports with final and differential diagnoses. One radiologist and one radiology resident independently reviewed cases without and with the reports created by ChatGPT‐4o at a 4‐week interval. Diagnostic accuracy was compared to the published gold‐standard histopathological diagnoses. Also, ChatGPT‐4o was asked to detect the orientation, sequence, and contrast usage. To compare the accuracies, exact McNemar's test and Wilcoxon signed‐rank test with Bonferroni correction was used for statistical analysis. Results Evaluation of radiology reports demonstrated an accuracy of final diagnosis from ChatGPT‐4o 33.33% (10/30), experienced radiologist 56.67% (17/30), and radiology resident 23.33% (7/30) ( p < 0.039 for experienced radiologist vs. ChatGPT‐4o and p < 0.375 for resident vs. ChatGPT‐4o). Accuracy rates for differential diagnosis were: ChatGPT‐4o 60.50% ± 21.30%, experienced radiologist 64.83% ± 17.24%, and radiology resident 42.50% ± 16.90% ( p = 0.242 for experienced radiologist vs. ChatGPT‐4o and p = 0.0004 for resident vs. ChatGPT‐4o). Diagnostic performance with the use of ChatGPT‐4o as a support tool demonstrated an accuracy of final diagnosis: experienced radiologist 56.67% (17/30) and radiology resident 30.00% (9/30) ( p = 1.000 and p = 0.500, respectively) and of differential diagnosis: experienced radiologist 66.50% ± 19.30% and radiology resident 49.50% ± 15.56% ( p = 0.500 and p = 0.013, respectively). Conclusions ChatGPT‐4o demonstrated limited autonomous and supportive diagnostic accuracy in complex neuroradiology cases. While its assistance did not improve the experienced radiologist's performance and produced only a modest, non‐significant improvement for the resident, a trend toward improved differential diagnosis generation was observed. These findings suggest that multimodal large language models may have limited supportive value, particularly for less experienced clinicians, and require validation in larger studies.
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.sponsoredbyTubitakEuN/A
dc.description.sponsorshipN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile45
dc.identifier.ScopusQuartileQ3
dc.identifier.WoSPercentile53.7
dc.identifier.WoSQuartileQ2
dc.identifier.doi10.1002/ird3.70085
dc.identifier.eissn2834-2879
dc.identifier.endpage359
dc.identifier.grantnoN/A
dc.identifier.issn2834-2860
dc.identifier.issue4
dc.identifier.scopus2-s2.0-105048463595
dc.identifier.startpage351
dc.identifier.urihttp://doi.org/10.1002/ird3.70085
dc.identifier.urihttps://hdl.handle.net/20.500.14288/35413
dc.identifier.volume4
dc.languageeng
dc.publisherWiley
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofIradiology
dc.relation.openaccessN/A
dc.subjectHealth sciences
dc.subjectMedicine
dc.subjectHealth informatics
dc.subjectRadiology
dc.subjectNuclear medicine and imaging
dc.titleAssessing chatgpt-4o's potential as a support tool in neuroradiology cases: a comparison of performance by using radiologic images
dc.typeJournal Article
dspace.entity.typePublication
relation.isOrgUnitOfPublicationf91d21f0-6b13-46ce-939a-db68e4c8d2ab
relation.isOrgUnitOfPublication.latestForDiscoveryf91d21f0-6b13-46ce-939a-db68e4c8d2ab
relation.isParentOrgUnitOfPublication055775c9-9efe-43ec-814f-f6d771fa6dee
relation.isParentOrgUnitOfPublication.latestForDiscovery055775c9-9efe-43ec-814f-f6d771fa6dee

Files