跳到主要导航 跳到搜索 跳到主要内容

Multi-MELO: Unified multimodal model editing with dynamic LoRA

  • East China Normal University

科研成果: 期刊稿件文章同行评审

摘要

Model editing aims to correct hallucinations or incorporate new knowledge into the pre-trained neural networks. Most previous researches focus on model editing with merely the textual modality, while editing for multimodal models is not well studied. Recent research investigates how to adapt language model editors to multimodal scenarios. However, these methods are limited to image-to-text tasks and similar model architectures. The text-to-image editing task remains unexplored, presenting significant challenges due to the diversity of complex network architectures. In this paper, we propose a unified multimodal model editing framework based on dynamic LoRA (Multi-MELO), which enables effective editing for various multimodal models by dynamically activating corresponding LoRA blocks that encode the related knowledge. We explore the framework for editing diverse multimodal models (i.e., BLIP-2, and latent diffusion model) on three downstream tasks, including image captioning, visual question answering and text-to-image generation. The experimental results show that Multi-MELO achieves superior editing performance compared to the recent state-of-the-art baselines, and meanwhile requires no extra training for additional modules.

源语言英语
文章编号126766
期刊Expert Systems with Applications
273
DOI
出版状态已出版 - 10 5月 2025

学术指纹

探究 'Multi-MELO: Unified multimodal model editing with dynamic LoRA' 的科研主题。它们共同构成独一无二的学术指纹。

引用此