Publication: Backdoor attacks in text classification: threats, methods, and emerging challenges
Program
KU-Authors
KU Authors
Co-Authors
Editor & Affiliation
Compiler & Affiliation
Translator
Other Contributor
Date
Language
eng
Type
Embargo Status
N/A
Journal Title
Journal ISSN
Volume Title
Alternative Title
Abstract
Backdoor attacks pose a serious threat to text classification models by embedding hidden malicious behaviors that activate when a specific trigger is present. They allow adversaries to manipulate model predictions while maintaining high accuracy on clean inputs, making detection difficult. In this chapter, we first present background on backdoor attacks and prominent attack methods in text classification. We then perform an experimental analysis demonstrating that attacks can achieve high success rates while preserving classification accuracy on clean inputs. We also discuss emerging issues, including clean-label attacks, attacks on large language models, and attacks and defenses in federated learning. By combining conceptual descriptions with empirical results, this chapter aims to provide an overview of backdoor attacks in text classification and potential future research directions.
Source
Publisher
Springer
Subject
Computer engineering, Computer science
Citation
Has Part
Source
Adversarial Example Detection and Mitigation Using Machine Learning
Book Series Title
Edition
DOI
10.1007/978-3-031-99447-0_4
item.page.datauri
Link
Rights
N/A
Copyrights Note
Creative Commons license
Except where otherwised noted, this item's license is described as N/A
