跳到主要导航 跳到搜索 跳到主要内容

Harvesting facts from textual web sources by constrained label propagation

  • Yafang Wang*
  • , Bin Yang
  • , Lizhen Qu
  • , Marc Spaniol
  • , Gerhard Weikum
  • *此作品的通讯作者
  • Max Planck Institute for Informatics

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

There have been major advances on automatically constructing large knowledge bases by extracting relational facts from Web and text sources. However, the world is dynamic: periodic events like sports competitions need to be interpreted with their respective timepoints, and facts such as coaching a sports team, holding political or business positions, and even marriages do not hold forever and should be augmented by their respective timespans. This paper addresses the problem of automatically harvesting temporal facts with such extended time-awareness. We employ pattern-based gathering techniques for fact candidates and construct a weighted pattern-candidate graph. Our key contribution is a system called PRAVDA based on a new kind of label propagation algorithm with a judiciously designed loss function, which iteratively processes the graph to label good temporal facts for a given set of target relations. Our experiments with online news and Wikipedia articles demonstrate the accuracy of this method.

源语言英语
主期刊名CIKM'11 - Proceedings of the 2011 ACM International Conference on Information and Knowledge Management
837-846
页数10
DOI
出版状态已出版 - 2011
已对外发布
活动20th ACM Conference on Information and Knowledge Management, CIKM'11 - Glasgow, 英国
期限: 24 10月 201128 10月 2011

出版系列

姓名International Conference on Information and Knowledge Management, Proceedings

会议

会议20th ACM Conference on Information and Knowledge Management, CIKM'11
国家/地区英国
Glasgow
时期24/10/1128/10/11

学术指纹

探究 'Harvesting facts from textual web sources by constrained label propagation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此