Defense Against Adversarial Attacks for Neural Representations of Text

dc.contributor.authorZhan, Huixin
dc.contributor.authorZhang, Kun
dc.contributor.authorChen, Zhong
dc.contributor.authorSheng, Victor
dc.date.accessioned2023-12-26T18:54:44Z
dc.date.available2023-12-26T18:54:44Z
dc.date.issued2024-01-03
dc.identifier.doihttps://doi.org/10.24251/HICSS.2024.912
dc.identifier.isbn978-0-9981331-7-1
dc.identifier.other306ecce4-2b6d-438b-b081-6f1c91d34c39
dc.identifier.urihttps://hdl.handle.net/10125/107298
dc.language.isoeng
dc.relation.ispartofProceedings of the 57th Hawaii International Conference on System Sciences
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 International
dc.rights.urihttps://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectMachine Learning and AI: Cybersecurity and Threat Hunting
dc.subjectadversary
dc.subjectattack
dc.subjectnatural language processing
dc.subjectprivacy-preserving
dc.subjecttext representations
dc.titleDefense Against Adversarial Attacks for Neural Representations of Text
dc.typeConference Paper
dc.type.dcmiText
dcterms.abstractIn this paper, we focus on defending against adversarial attacks for privacy-preserving Natural Language Processing (NLP) under a model partitioning scenario, where the model splits into a local, on-device part and a remote, cloud-based part. Model partitioning improves the scalability and protects the privacy of inputs into the model. However, we argue that privacy protection breaks during inference with model partitioning. In this paper, an adversary eavesdrops on the hidden representations output from the local devices and tries to use the representations to obtain private information from the input text. We study two types of adversarial attacks, i.e., adversarial classification and adversarial generation. Based on these two attack models, we correspondingly propose two defenses: defending the adversarial classification (DAC) and defending the adversarial generation (DAG). Specifically, the DAC and DAG approaches are both bilevel optimization-based defense methods. Both methods optimally modify a subpopulation of the neural representations that are subject to maximally decreasing the adversary’s ability. The representations trained with this bilevel optimization protect sensitive information from the adversary attack while maintaining their utility for downstream tasks. Our experiments show that both DAC and DAG approaches improve the performance of the main text classifier and achieve even higher privacy of neural representations compared with the current state-of-the-art methods.
dcterms.extent10 pages
prism.startingpage7592

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
0741.pdf
Size:
3.11 MB
Format:
Adobe Portable Document Format