Publication:
Backdoor attacks in text classification: threats, methods, and emerging challenges

dc.contributor.departmentGraduate School of Sciences and Engineering
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.kuauthorEralp, Egehan
dc.contributor.kuauthorGürsoy, Mehmet Emre
dc.contributor.schoolcollegeinstituteGRADUATE SCHOOL OF SCIENCES AND ENGINEERING
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.date.accessioned2026-08-14T11:20:27Z
dc.date.issued2026
dc.description.abstractBackdoor attacks pose a serious threat to text classification models by embedding hidden malicious behaviors that activate when a specific trigger is present. They allow adversaries to manipulate model predictions while maintaining high accuracy on clean inputs, making detection difficult. In this chapter, we first present background on backdoor attacks and prominent attack methods in text classification. We then perform an experimental analysis demonstrating that attacks can achieve high success rates while preserving classification accuracy on clean inputs. We also discuss emerging issues, including clean-label attacks, attacks on large language models, and attacks and defenses in federated learning. By combining conceptual descriptions with empirical results, this chapter aims to provide an overview of backdoor attacks in text classification and potential future research directions.
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipThis research was supported by the Scientific and Technological Research Council of Türkiye (TUBITAK) under Grant Number 125E059. The authors thank TUBITAK for their support.
dc.description.versionPublished Version
dc.identifier.ScopusPercentileN/A
dc.identifier.ScopusQuartileN/A
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.1007/978-3-031-99447-0_4
dc.identifier.embargoN/A
dc.identifier.endpage63
dc.identifier.grantno125E059
dc.identifier.isbn9783031994463
dc.identifier.scopus2-s2.0-105030223858
dc.identifier.startpage51
dc.identifier.urihttp://doi.org/10.1007/978-3-031-99447-0_4
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34315
dc.keywordsAI security
dc.keywordsBackdoor attacks
dc.keywordsNatural language processing (NLP)
dc.keywordsText classification
dc.languageeng
dc.publisherSpringer
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofAdversarial Example Detection and Mitigation Using Machine Learning
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectComputer engineering
dc.subjectComputer science
dc.titleBackdoor attacks in text classification: threats, methods, and emerging challenges
dc.typeBook Chapter
dspace.entity.typePublication
relation.isOrgUnitOfPublication3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isParentOrgUnitOfPublication434c9663-2b11-4e66-9399-c863e2ebae43
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery434c9663-2b11-4e66-9399-c863e2ebae43

Files