Publication:
Backdoor attacks in text classification: threats, methods, and emerging challenges

Placeholder

School / College / Institute

Organizational Unit

Program

KU Authors

Co-Authors

Editor & Affiliation

Compiler & Affiliation

Translator

Other Contributor

Date

Language

eng

Embargo Status

N/A

Journal Title

Journal ISSN

Volume Title

Alternative Title

Abstract

Backdoor attacks pose a serious threat to text classification models by embedding hidden malicious behaviors that activate when a specific trigger is present. They allow adversaries to manipulate model predictions while maintaining high accuracy on clean inputs, making detection difficult. In this chapter, we first present background on backdoor attacks and prominent attack methods in text classification. We then perform an experimental analysis demonstrating that attacks can achieve high success rates while preserving classification accuracy on clean inputs. We also discuss emerging issues, including clean-label attacks, attacks on large language models, and attacks and defenses in federated learning. By combining conceptual descriptions with empirical results, this chapter aims to provide an overview of backdoor attacks in text classification and potential future research directions.

Source

Publisher

Springer

Subject

Computer engineering, Computer science

Citation

Has Part

Source

Adversarial Example Detection and Mitigation Using Machine Learning

Book Series Title

Edition

DOI

10.1007/978-3-031-99447-0_4

item.page.datauri

Link

Rights

N/A

Copyrights Note

Creative Commons license

Except where otherwised noted, this item's license is described as N/A

Endorsement

Review

Supplemented By

Referenced By

Related Goal

0

Views

0

Downloads

View PlumX Details