跳到主要导航 跳到搜索 跳到主要内容

SGFL-Attack: A Similarity-Guidance Strategy for Hard-Label Textual Adversarial Attack Based on Feedback Learning

  • Panjia Qiu
  • , Guanghao Zhou
  • , Mingyuan Fan
  • , Cen Chen*
  • , Yaliang Li
  • , Wenming Zhou
  • *此作品的通讯作者
  • East China Normal University
  • Zhejiang University
  • Alibaba Group Holding Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Hard-label black-box textual adversarial attack presents a challenging task where only the predictions of the victim model are available. Moreover, several constraints further complicate the task of launching such attacks, including the inherent discrete and non-differentiable nature of text data and the need to introduce subtle perturbations that remain imperceptible to humans while preserving semantic similarity. Despite the considerable research efforts dedicated to this problem, existing methods still suffer from several limitations. For example, algorithms based on complex heuristic searches necessitate extensive querying, rendering them computationally expensive. The introduction of continuous gradient strategies into discrete text spaces often leads to estimation errors. Meanwhile, geometry-based strategies are prone to falling into local optima. To address these limitations, in this paper, we introduce SGFL-Attack, a novel approach that leverages a <u>S</u>imilarity-<u>G</u>uidance strategy based on <u>F</u>eedback <u>L</u>earning for hard-label textual adversarial attack, with limited query budget. Specifically, the proposed SGFL-Attack utilizes word embedding vectors to assess the importance of words and positions in text sequences, and employs a feedback learning mechanism to determine reward or punishment based on changes in predicted labels caused by replacing words. In each iteration, SGFL-Attack guides the search based on knowledge acquired from the feedback learning mechanism, generating more similar samples while maintaining low perturbations. Moreover, to reduce the query budget, we incorporate local hash mapping to avoid redundant queries during the search process. Extensive experiments on seven widely used datasets show that the proposed SGFL-Attack method significantly outperforms state-of-the-art baselines and defenses over multiple language models.

源语言英语
主期刊名CIKM 2024 - Proceedings of the 33rd ACM International Conference on Information and Knowledge Management
出版商Association for Computing Machinery
1920-1929
页数10
ISBN(电子版)9798400704369
DOI
出版状态已出版 - 21 10月 2024
活动33rd ACM International Conference on Information and Knowledge Management, CIKM 2024 - Boise, 美国
期限: 21 10月 202425 10月 2024

出版系列

姓名International Conference on Information and Knowledge Management, Proceedings
ISSN(印刷版)2155-0751

会议

会议33rd ACM International Conference on Information and Knowledge Management, CIKM 2024
国家/地区美国
Boise
时期21/10/2425/10/24

指纹

探究 'SGFL-Attack: A Similarity-Guidance Strategy for Hard-Label Textual Adversarial Attack Based on Feedback Learning' 的科研主题。它们共同构成独一无二的指纹。

引用此