Abstract
The current dynamic scene graph generation (DSGG) methods rely on dense annotations, which are costly and have significant limitations in fine-grained relationship prediction. Although few-shot learning can achieve rapid adaptation with a small number of annotated samples, the diversity of object-predicate combinations in video scenes leads to significant feature heterogeneity, while the temporal complexity of dynamic scenes increases the difficulty of predicate prediction. Therefore, this paper pioneers a DSGG method for few-shot learning of multimodal enhanced dynamic prototype learning (MEDP). Specifically, the method comprises two key modules: (1) Multimodal Feature Enhancement (MFE), which generates textual descriptions of the scene via a visual language model and fuses them with frame-level visual features to enhance the representation of dynamic scene graphs. Additionally, MEDP learning uses temporal modeling to capture dynamic changes and long temporal dependencies in videos to enhance the capacity of models to understand the temporal context. (2) Dynamic Prototype Matching (DPM) achieves predicate matching in a few-shot setting by dynamically modeling each prototype based on object category and positional encoding; it computes relational instance prototypes across the textual, frame and serial levels, combining them via a weighted fusion of predicate predictions. Experiments demonstrated that the method exhibited remarkable generalizability and stability in few-shot DSGG.
| Original language | English |
|---|---|
| Article number | 116323 |
| Journal | Knowledge-Based Systems |
| Volume | 348 |
| DOIs | |
| State | Published - 3 Aug 2026 |
Keywords
- Dynamic prototype matching
- Dynamic scene graph generation
- Few-shot learning
- Multimodal feature enhancement
Fingerprint
Dive into the research topics of 'MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver