跳到主要导航 跳到搜索 跳到主要内容

Black-box Prompt Tuning for Vision-Language Model as a Service

  • East China Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

In the scenario of Model-as-a-Service (MaaS), pre-trained models are usually released as inference APIs. Users are allowed to query those models with manually crafted prompts. Without accessing the network structure and gradient information, it's tricky to perform continuous prompt tuning on MaaS, especially for vision-language models (VLMs) considering cross-modal interaction. In this paper, we propose a black-box prompt tuning framework for VLMs to learn task-relevant prompts without back-propagation. In particular, the vision and language prompts are jointly optimized in the intrinsic parameter subspace with various evolution strategies. Different prompt variants are also explored to enhance the cross-model interaction. Experimental results show that our proposed black-box prompt tuning framework outperforms both hand-crafted prompt engineering and gradient-based prompt learning methods, which serves as evidence of its capability to train task-relevant prompts in a derivative-free manner.

源语言英语
主期刊名Proceedings of the 32nd International Joint Conference on Artificial Intelligence, IJCAI 2023
编辑Edith Elkind
出版商International Joint Conferences on Artificial Intelligence
1686-1694
页数9
ISBN(电子版)9781956792034
DOI
出版状态已出版 - 2023
活动32nd International Joint Conference on Artificial Intelligence, IJCAI 2023 - Macao, 中国
期限: 19 8月 202325 8月 2023

出版系列

姓名IJCAI International Joint Conference on Artificial Intelligence
2023-August
ISSN(印刷版)1045-0823

会议

会议32nd International Joint Conference on Artificial Intelligence, IJCAI 2023
国家/地区中国
Macao
时期19/08/2325/08/23

指纹

探究 'Black-box Prompt Tuning for Vision-Language Model as a Service' 的科研主题。它们共同构成独一无二的指纹。

引用此