Publication:
Strategic sample selection for improved clean-label backdoor attacks in text classification

dc.conference.dateSEP 22-24, 2025
dc.contributor.coauthorKirci, O. A.
dc.contributor.coauthorGursoy, M. E.
dc.date.accessioned2026-08-14T11:27:00Z
dc.date.issued2026
dc.description.abstractBackdoor attacks pose a significant threat to the integrity of text classification models used in natural language processing. While several dirty-label attacks that achieve high attack success rates (ASR) have been proposed, clean-label attacks are inherently more difficult. In this paper, we propose three sample selection strategies to improve attack effectiveness in clean-label scenarios: Minimum, Above50, and Below50. Our strategies identify those samples which the model predicts incorrectly or with low confidence, and by injecting backdoor triggers into such samples, we aim to induce a stronger association between the trigger patterns and the attacker-desired target label. We apply our methods to clean-label variants of four canonical backdoor attacks (InsertSent, WordInj, StyleBkd, SynBkd) and evaluate them on three datasets (IMDB, SST2, HateSpeech) and four model types (LSTM, BERT, DistilBERT, RoBERTa). Results show that the proposed strategies, particularly the Minimum strategy, significantly improve the ASR over random sample selection with little or no degradation in the model’s clean accuracy. Furthermore, clean-label attacks enhanced by our strategies outperform BITE, a state of the art clean-label attack method, in many configurations
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipThis study was supported by The Scientific and Technological Research Council of Turkiye (TUBITAK) under grant number 125E059 and the BAGEP Outstanding Young Scientist Award. The authors thank TUBITAK and the Science Academy for their support
dc.description.versionPublished Version
dc.identifier.ScopusPercentile53
dc.identifier.ScopusQuartileQ2
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.1007/978-3-032-16092-8_22
dc.identifier.eissn1611-3349
dc.identifier.embargoN/A
dc.identifier.endpage415
dc.identifier.grantno125E059
dc.identifier.issn0302-9743
dc.identifier.scopus2-s2.0-105039701327
dc.identifier.startpage397
dc.identifier.urihttp://doi.org/10.1007/978-3-032-16092-8_22
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34641
dc.keywordsAdversarial machine learning
dc.keywordsAI security
dc.keywordsBackdoor attacks
dc.keywordsLanguage models
dc.keywordsNatural language processing
dc.keywordsText classification
dc.languageeng
dc.publisherSpringer
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofLecture Notes in Computer Science
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectPhysical sciences
dc.subjectComputer science
dc.subjectArtificial intelligence
dc.subjectInformation systems
dc.titleStrategic sample selection for improved clean-label backdoor attacks in text classification
dc.typeConference Proceeding
dspace.entity.typePublication

Files