Publication: Backdoor attacks in text classification: threats, methods, and emerging challenges
| dc.contributor.department | Graduate School of Sciences and Engineering | |
| dc.contributor.department | Department of Computer Engineering | |
| dc.contributor.kuauthor | Eralp, Egehan | |
| dc.contributor.kuauthor | Gürsoy, Mehmet Emre | |
| dc.contributor.schoolcollegeinstitute | GRADUATE SCHOOL OF SCIENCES AND ENGINEERING | |
| dc.contributor.schoolcollegeinstitute | College of Engineering | |
| dc.date.accessioned | 2026-08-14T11:20:27Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Backdoor attacks pose a serious threat to text classification models by embedding hidden malicious behaviors that activate when a specific trigger is present. They allow adversaries to manipulate model predictions while maintaining high accuracy on clean inputs, making detection difficult. In this chapter, we first present background on backdoor attacks and prominent attack methods in text classification. We then perform an experimental analysis demonstrating that attacks can achieve high success rates while preserving classification accuracy on clean inputs. We also discuss emerging issues, including clean-label attacks, attacks on large language models, and attacks and defenses in federated learning. By combining conceptual descriptions with empirical results, this chapter aims to provide an overview of backdoor attacks in text classification and potential future research directions. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | TÜBİTAK | |
| dc.description.sponsorship | This research was supported by the Scientific and Technological Research Council of Türkiye (TUBITAK) under Grant Number 125E059. The authors thank TUBITAK for their support. | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | N/A | |
| dc.identifier.ScopusQuartile | N/A | |
| dc.identifier.WoSPercentile | N/A | |
| dc.identifier.WoSQuartile | N/A | |
| dc.identifier.doi | 10.1007/978-3-031-99447-0_4 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.endpage | 63 | |
| dc.identifier.grantno | 125E059 | |
| dc.identifier.isbn | 9783031994463 | |
| dc.identifier.scopus | 2-s2.0-105030223858 | |
| dc.identifier.startpage | 51 | |
| dc.identifier.uri | http://doi.org/10.1007/978-3-031-99447-0_4 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34315 | |
| dc.keywords | AI security | |
| dc.keywords | Backdoor attacks | |
| dc.keywords | Natural language processing (NLP) | |
| dc.keywords | Text classification | |
| dc.language | eng | |
| dc.publisher | Springer | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | Adversarial Example Detection and Mitigation Using Machine Learning | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Computer engineering | |
| dc.subject | Computer science | |
| dc.title | Backdoor attacks in text classification: threats, methods, and emerging challenges | |
| dc.type | Book Chapter | |
| dspace.entity.type | Publication | |
| relation.isOrgUnitOfPublication | 3fc31c89-e803-4eb1-af6b-6258bc42c3d8 | |
| relation.isOrgUnitOfPublication | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isOrgUnitOfPublication.latestForDiscovery | 3fc31c89-e803-4eb1-af6b-6258bc42c3d8 | |
| relation.isParentOrgUnitOfPublication | 434c9663-2b11-4e66-9399-c863e2ebae43 | |
| relation.isParentOrgUnitOfPublication | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 | |
| relation.isParentOrgUnitOfPublication.latestForDiscovery | 434c9663-2b11-4e66-9399-c863e2ebae43 |
