跳到主要导航 跳到搜索 跳到主要内容

MM-Path: Multi-modal, Multi-granularity Path Representation Learning

  • Ronghui Xu
  • , Hanyin Cheng
  • , Chenjuan Guo
  • , Hongfan Gao
  • , Jilin Hu
  • , Sean Bin Yang
  • , Bin Yang*
  • *此作品的通讯作者
  • East China Normal University
  • KLATASDS-MOE
  • Chongqing University of Posts and Telecommunications

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performance, they predominantly focus on the topological structures from single modality data, i.e., road networks, overlooking the geometric and contextual features associated with path-related images, e.g., remote sensing images. Similar to human understanding, integrating information from multiple modalities can provide a more comprehensive view, enhancing both representation accuracy and generalization. However, variations in information granularity impede the semantic alignment of road network-based paths (road paths) and image-based paths (image paths), while the heterogeneity of multi-modal data poses substantial challenges for effective fusion and utilization. In this paper, we propose a novel Multi-modal, Multi-granularity Path Representation Learning Framework (MM-Path), which can learn a generic path representation by integrating modalities from both road paths and image paths. To enhance the alignment of multi-modal data, we develop a multi-granularity alignment strategy that systematically associates nodes, road sub-paths, and road paths with their corresponding image patches, ensuring the synchronization of both detailed local information and broader global contexts. To address the heterogeneity of multi-modal data effectively, we introduce a graph-based cross-modal residual fusion component designed to comprehensively fuse information across different modalities and granularities. Finally, we conduct extensive experiments on two large-scale real-world datasets under two downstream tasks, validating the effectiveness of the proposed MM-Path.

源语言英语
主期刊名KDD 2025 - Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
出版商Association for Computing Machinery
1703-1714
页数12
ISBN(电子版)9798400712456
DOI
出版状态已出版 - 20 7月 2025
活动31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025 - Toronto, 加拿大
期限: 3 8月 20257 8月 2025

出版系列

姓名Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
1
ISSN(印刷版)2154-817X

会议

会议31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025
国家/地区加拿大
Toronto
时期3/08/257/08/25

指纹

探究 'MM-Path: Multi-modal, Multi-granularity Path Representation Learning' 的科研主题。它们共同构成独一无二的指纹。

引用此