跳到主要导航 跳到搜索 跳到主要内容

MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation

  • East China Normal University
  • CAS - Institute of Automation

科研成果: 期刊稿件文章同行评审

摘要

The current dynamic scene graph generation (DSGG) methods rely on dense annotations, which are costly and have significant limitations in fine-grained relationship prediction. Although few-shot learning can achieve rapid adaptation with a small number of annotated samples, the diversity of object-predicate combinations in video scenes leads to significant feature heterogeneity, while the temporal complexity of dynamic scenes increases the difficulty of predicate prediction. Therefore, this paper pioneers a DSGG method for few-shot learning of multimodal enhanced dynamic prototype learning (MEDP). Specifically, the method comprises two key modules: (1) Multimodal Feature Enhancement (MFE), which generates textual descriptions of the scene via a visual language model and fuses them with frame-level visual features to enhance the representation of dynamic scene graphs. Additionally, MEDP learning uses temporal modeling to capture dynamic changes and long temporal dependencies in videos to enhance the capacity of models to understand the temporal context. (2) Dynamic Prototype Matching (DPM) achieves predicate matching in a few-shot setting by dynamically modeling each prototype based on object category and positional encoding; it computes relational instance prototypes across the textual, frame and serial levels, combining them via a weighted fusion of predicate predictions. Experiments demonstrated that the method exhibited remarkable generalizability and stability in few-shot DSGG.

源语言英语
期刊论文编号116323
期刊Knowledge-Based Systems
348
DOI
出版状态已出版 - 3 8月 2026

学术指纹

探究 'MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此