Skip to main navigation Skip to search Skip to main content

MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation

  • Xuejiao Wang
  • , Ziheng Huang
  • , Weiliang Meng
  • , Changbo Wang*
  • , Gaoqi He*
  • *Corresponding author for this work
  • East China Normal University
  • CAS - Institute of Automation

Research output: Contribution to journalArticlepeer-review

Abstract

The current dynamic scene graph generation (DSGG) methods rely on dense annotations, which are costly and have significant limitations in fine-grained relationship prediction. Although few-shot learning can achieve rapid adaptation with a small number of annotated samples, the diversity of object-predicate combinations in video scenes leads to significant feature heterogeneity, while the temporal complexity of dynamic scenes increases the difficulty of predicate prediction. Therefore, this paper pioneers a DSGG method for few-shot learning of multimodal enhanced dynamic prototype learning (MEDP). Specifically, the method comprises two key modules: (1) Multimodal Feature Enhancement (MFE), which generates textual descriptions of the scene via a visual language model and fuses them with frame-level visual features to enhance the representation of dynamic scene graphs. Additionally, MEDP learning uses temporal modeling to capture dynamic changes and long temporal dependencies in videos to enhance the capacity of models to understand the temporal context. (2) Dynamic Prototype Matching (DPM) achieves predicate matching in a few-shot setting by dynamically modeling each prototype based on object category and positional encoding; it computes relational instance prototypes across the textual, frame and serial levels, combining them via a weighted fusion of predicate predictions. Experiments demonstrated that the method exhibited remarkable generalizability and stability in few-shot DSGG.

Original languageEnglish
Article number116323
JournalKnowledge-Based Systems
Volume348
DOIs
StatePublished - 3 Aug 2026

Keywords

  • Dynamic prototype matching
  • Dynamic scene graph generation
  • Few-shot learning
  • Multimodal feature enhancement

Fingerprint

Dive into the research topics of 'MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation'. Together they form a unique fingerprint.

Cite this