TY - JOUR
T1 - MEDP
T2 - Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation
AU - Wang, Xuejiao
AU - Huang, Ziheng
AU - Meng, Weiliang
AU - Wang, Changbo
AU - He, Gaoqi
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/8/3
Y1 - 2026/8/3
N2 - The current dynamic scene graph generation (DSGG) methods rely on dense annotations, which are costly and have significant limitations in fine-grained relationship prediction. Although few-shot learning can achieve rapid adaptation with a small number of annotated samples, the diversity of object-predicate combinations in video scenes leads to significant feature heterogeneity, while the temporal complexity of dynamic scenes increases the difficulty of predicate prediction. Therefore, this paper pioneers a DSGG method for few-shot learning of multimodal enhanced dynamic prototype learning (MEDP). Specifically, the method comprises two key modules: (1) Multimodal Feature Enhancement (MFE), which generates textual descriptions of the scene via a visual language model and fuses them with frame-level visual features to enhance the representation of dynamic scene graphs. Additionally, MEDP learning uses temporal modeling to capture dynamic changes and long temporal dependencies in videos to enhance the capacity of models to understand the temporal context. (2) Dynamic Prototype Matching (DPM) achieves predicate matching in a few-shot setting by dynamically modeling each prototype based on object category and positional encoding; it computes relational instance prototypes across the textual, frame and serial levels, combining them via a weighted fusion of predicate predictions. Experiments demonstrated that the method exhibited remarkable generalizability and stability in few-shot DSGG.
AB - The current dynamic scene graph generation (DSGG) methods rely on dense annotations, which are costly and have significant limitations in fine-grained relationship prediction. Although few-shot learning can achieve rapid adaptation with a small number of annotated samples, the diversity of object-predicate combinations in video scenes leads to significant feature heterogeneity, while the temporal complexity of dynamic scenes increases the difficulty of predicate prediction. Therefore, this paper pioneers a DSGG method for few-shot learning of multimodal enhanced dynamic prototype learning (MEDP). Specifically, the method comprises two key modules: (1) Multimodal Feature Enhancement (MFE), which generates textual descriptions of the scene via a visual language model and fuses them with frame-level visual features to enhance the representation of dynamic scene graphs. Additionally, MEDP learning uses temporal modeling to capture dynamic changes and long temporal dependencies in videos to enhance the capacity of models to understand the temporal context. (2) Dynamic Prototype Matching (DPM) achieves predicate matching in a few-shot setting by dynamically modeling each prototype based on object category and positional encoding; it computes relational instance prototypes across the textual, frame and serial levels, combining them via a weighted fusion of predicate predictions. Experiments demonstrated that the method exhibited remarkable generalizability and stability in few-shot DSGG.
KW - Dynamic prototype matching
KW - Dynamic scene graph generation
KW - Few-shot learning
KW - Multimodal feature enhancement
UR - https://www.scopus.com/pages/publications/105041116214
U2 - 10.1016/j.knosys.2026.116323
DO - 10.1016/j.knosys.2026.116323
M3 - 文章
AN - SCOPUS:105041116214
SN - 0950-7051
VL - 348
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116323
ER -