Publication:
DeVisE: towards the behavioral testing of medical large language models

dc.conference.dateMAR 24-29, 2026
dc.conference.locationRabat, Morocco
dc.contributor.coauthorTagliabue, C. Z.
dc.contributor.coauthorBoll, H. O.
dc.contributor.coauthorErdem, E.
dc.contributor.coauthorCalixto, I.
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.kuauthorErdem, Aykut
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.date.accessioned2026-08-14T11:21:42Z
dc.date.issued2026
dc.description.abstractLarge language models (LLMs) are increasingly applied in clinical decision support, yet current evaluations rarely reveal whether their outputs reflect genuine medical reasoning or superficial correlations.We introduce DeVisE (Demographics and Vital signs Evaluation), a behavioral testing framework that probes finegrained clinical understanding through controlled counterfactuals.Using intensive care unit (ICU) discharge notes from MIMIC-IV, we construct both raw (real-world) and templatebased (synthetic) variants with single-variable perturbations in demographic (age, gender, ethnicity) and vital sign attributes.We evaluate eight LLMs, spanning general-purpose and medical variants, under zero-shot setting.Model behavior is analyzed through (1) inputlevel sensitivity, capturing how counterfactuals alter perplexity, and (2) downstream reasoning, measuring their effect on predicted ICU lengthof-stay and mortality.Overall, our results show that standard task metrics obscure clinically relevant differences in model behavior, with models differing substantially in how consistently and proportionally they adjust predictions to counterfactual perturbations. 1 How do counterfactuals change the probability distribution of downstream tasks?How do counterfactuals change how likely the patient is?OpenBioLLM 70B LLaMA-3.3-Instruct70B DeepSeek-R1-Distill 70B Qwen-2.5-Instruct72B GPT-OSS 120B GPT-4.1-mini?
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipHOB and IC are funded by the project CaRe-NLP with file number NGF.1607.22.014 of the research programme AiNed Fellowship Grants which is (partly) financed by the Dutch Research Council (NWO). EE is funded by the project TUBITAK 2247-A National Outstanding Researchers Program Award No. 123C542.
dc.description.versionPublished Version
dc.identifier.ScopusPercentileN/A
dc.identifier.ScopusQuartileN/A
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.18653/v1/2026.findings-eacl.338
dc.identifier.embargoN/A
dc.identifier.endpage6441
dc.identifier.grantno123C542
dc.identifier.isbn9798891763869
dc.identifier.scopus2-s2.0-105039133494
dc.identifier.startpage6427
dc.identifier.urihttp://doi.org/10.18653/v1/2026.findings-eacl.338
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34387
dc.keywordsBehavioral analysis
dc.keywordsNatural language
dc.keywordsLanguage model
dc.keywordsTest (biology)
dc.languageeng
dc.publisherAssociation for Computational Linguistics
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofFindings of the Association for Computational Linguistics: Eacl 2026
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectMedicine
dc.subjectHealth informatics
dc.titleDeVisE: towards the behavioral testing of medical large language models
dc.typeConference Proceeding
dspace.entity.typePublication
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery8e756b23-2d4a-4ce8-b1b3-62c794a8c164

Files