跳到主要导航 跳到搜索 跳到主要内容

Perception-Guided Jailbreak Against Text-to-Image Models

  • Yihao Huang
  • , Le Liang*
  • , Tianlin Li
  • , Xiaojun Jia*
  • , Run Wang
  • , Weikai Miao
  • , Geguang Pu
  • , Yang Liu
  • *此作品的通讯作者
  • Nanyang Technological University
  • East China Normal University
  • Ministry of Education of the People's Republic of China
  • Wuhan University
  • Shanghai Trusted Industrial Control Platform Co., Ltd.

科研成果: 期刊稿件会议文章同行评审

摘要

In recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ. Warning: This paper contains NSFW and disturbing imagery, including adult, violent, and illegal-related contentious content. We have masked images deemed unsafe. However, reader discretion is advised.

源语言英语
页(从-至)26238-26247
页数10
期刊Proceedings of the AAAI Conference on Artificial Intelligence
39
25
DOI
出版状态已出版 - 11 4月 2025
活动39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, 美国
期限: 25 2月 20254 3月 2025

指纹

探究 'Perception-Guided Jailbreak Against Text-to-Image Models' 的科研主题。它们共同构成独一无二的指纹。

引用此