Publication: DeVisE: towards the behavioral testing of medical large language models
| dc.conference.date | MAR 24-29, 2026 | |
| dc.conference.location | Rabat, Morocco | |
| dc.contributor.coauthor | Tagliabue, C. Z. | |
| dc.contributor.coauthor | Boll, H. O. | |
| dc.contributor.coauthor | Erdem, E. | |
| dc.contributor.coauthor | Calixto, I. | |
| dc.contributor.department | Department of Computer Engineering | |
| dc.contributor.kuauthor | Erdem, Aykut | |
| dc.contributor.schoolcollegeinstitute | College of Engineering | |
| dc.date.accessioned | 2026-08-14T11:21:42Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Large language models (LLMs) are increasingly applied in clinical decision support, yet current evaluations rarely reveal whether their outputs reflect genuine medical reasoning or superficial correlations.We introduce DeVisE (Demographics and Vital signs Evaluation), a behavioral testing framework that probes finegrained clinical understanding through controlled counterfactuals.Using intensive care unit (ICU) discharge notes from MIMIC-IV, we construct both raw (real-world) and templatebased (synthetic) variants with single-variable perturbations in demographic (age, gender, ethnicity) and vital sign attributes.We evaluate eight LLMs, spanning general-purpose and medical variants, under zero-shot setting.Model behavior is analyzed through (1) inputlevel sensitivity, capturing how counterfactuals alter perplexity, and (2) downstream reasoning, measuring their effect on predicted ICU lengthof-stay and mortality.Overall, our results show that standard task metrics obscure clinically relevant differences in model behavior, with models differing substantially in how consistently and proportionally they adjust predictions to counterfactual perturbations. 1 How do counterfactuals change the probability distribution of downstream tasks?How do counterfactuals change how likely the patient is?OpenBioLLM 70B LLaMA-3.3-Instruct70B DeepSeek-R1-Distill 70B Qwen-2.5-Instruct72B GPT-OSS 120B GPT-4.1-mini? | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | TÜBİTAK | |
| dc.description.sponsorship | HOB and IC are funded by the project CaRe-NLP with file number NGF.1607.22.014 of the research programme AiNed Fellowship Grants which is (partly) financed by the Dutch Research Council (NWO). EE is funded by the project TUBITAK 2247-A National Outstanding Researchers Program Award No. 123C542. | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | N/A | |
| dc.identifier.ScopusQuartile | N/A | |
| dc.identifier.WoSPercentile | N/A | |
| dc.identifier.WoSQuartile | N/A | |
| dc.identifier.doi | 10.18653/v1/2026.findings-eacl.338 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.endpage | 6441 | |
| dc.identifier.grantno | 123C542 | |
| dc.identifier.isbn | 9798891763869 | |
| dc.identifier.scopus | 2-s2.0-105039133494 | |
| dc.identifier.startpage | 6427 | |
| dc.identifier.uri | http://doi.org/10.18653/v1/2026.findings-eacl.338 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34387 | |
| dc.keywords | Behavioral analysis | |
| dc.keywords | Natural language | |
| dc.keywords | Language model | |
| dc.keywords | Test (biology) | |
| dc.language | eng | |
| dc.publisher | Association for Computational Linguistics | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | Findings of the Association for Computational Linguistics: Eacl 2026 | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Medicine | |
| dc.subject | Health informatics | |
| dc.title | DeVisE: towards the behavioral testing of medical large language models | |
| dc.type | Conference Proceeding | |
| dspace.entity.type | Publication | |
| relation.isOrgUnitOfPublication | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isOrgUnitOfPublication.latestForDiscovery | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isParentOrgUnitOfPublication | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 | |
| relation.isParentOrgUnitOfPublication.latestForDiscovery | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 |
