Publication:
Win-k: improved membership inference attacks on small language models

dc.conference.dateSEP 22-24, 2025
dc.contributor.coauthorArkhmammadova, R.
dc.contributor.coauthorTamar, H. M.
dc.contributor.coauthorGursoy, M. E.
dc.contributor.departmentGraduate School of Sciences and Engineering
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.kuauthorArkhmammadova, Roya
dc.contributor.kuauthorGürsoy, Mehmet Emre
dc.contributor.kuauthorTamar, Hosein Madadi
dc.contributor.schoolcollegeinstituteGRADUATE SCHOOL OF SCIENCES AND ENGINEERING
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.date.accessioned2026-08-14T11:27:59Z
dc.date.issued2026
dc.description.abstractSmall language models (SLMs) are increasingly valued for their efficiency and deployability in resource-constrained environments, making them useful for on-device, privacy-sensitive, and edge computing applications. On the other hand, membership inference attacks (MIAs), which aim to determine whether a given sample was used in a model’s training, are an important threat with serious privacy and intellectual property implications. In this paper, we study MIAs on SLMs. Although MIAs were shown to be effective on large language models (LLMs), they are relatively less studied on emerging SLMs, and furthermore, their effectiveness decreases as models get smaller. Motivated by this finding, we propose a new MIA called win-k, which builds on top of a state-of-the-art attack (min-k). We experimentally evaluate win-k by comparing it with five existing MIAs using three datasets and eight SLMs. Results show that win-k outperforms existing MIAs in terms of AUROC, TPR @ 1% FPR, and FPR @ 99% TPR metrics, especially on smaller models.
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipThis study was supported by The Scientific and Technological Research Council of Turkiye (TUBITAK) under grants numbered 123E179 and 125E059. The authors thank TUBITAK for their support
dc.description.versionPublished Version
dc.identifier.ScopusPercentile53
dc.identifier.ScopusQuartileQ2
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.1007/978-3-032-16089-8_5
dc.identifier.eissn1611-3349
dc.identifier.embargoN/A
dc.identifier.endpage78
dc.identifier.grantno123E179
dc.identifier.grantno125E059
dc.identifier.isbn9783032160881
dc.identifier.issn0302-9743
dc.identifier.scopus2-s2.0-105039336485
dc.identifier.startpage66
dc.identifier.urihttp://doi.org/10.1007/978-3-032-16089-8_5
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34700
dc.keywordsAI security
dc.keywordsMembership inference attacks
dc.keywordsPrivacy
dc.keywordsResponsible AI
dc.keywordsSmall language models
dc.languageeng
dc.publisherSpringer
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofLecture Notes in Computer Science
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectComputer engineering
dc.subjectComputer science
dc.subjectArtificial intelligence
dc.titleWin-k: improved membership inference attacks on small language models
dc.typeConference Proceeding
dspace.entity.typePublication
relation.isOrgUnitOfPublication3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isParentOrgUnitOfPublication434c9663-2b11-4e66-9399-c863e2ebae43
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery434c9663-2b11-4e66-9399-c863e2ebae43

Files