跳到主要导航 跳到搜索 跳到主要内容

Building a Web Thesaurus from Web Link Structure

  • Zheng Chen*
  • , Shengping Liu
  • , Liu Wenyin
  • , Geguang Pu
  • , Wei Ying Ma
  • *此作品的通讯作者
  • Microsoft USA
  • Peking University
  • City University of Hong Kong

科研成果: 期刊稿件会议文章同行评审

摘要

Thesaurus has been widely used in many applications, including information retrieval, natural language processing, and question answering. In this paper, we propose a novel approach to automatically constructing a domain-specific thesaurus from the Web using link structure information. The proposed approach is able to identify new terms and reflect the latest relationship between terms as the Web evolves. First, a set of high quality and representative websites of a specific domain is selected. After filtering out navigational links, link analysis is applied to each website to obtain its content structure. Finally, the thesaurus is constructed by merging the content structures of the selected websites. The experimental results on automatic query expansion based on our constructed thesaurus show 20% improvement in search precision compared to the baseline.

源语言英语
页(从-至)48-55
页数8
期刊SIGIR Forum (ACM Special Interest Group on Information Retrieval)
SPEC. ISS.
DOI
出版状态已出版 - 2003
已对外发布
活动Proceedings of the Twenty-Sixth Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2003 - Toronto, Ont., 加拿大
期限: 28 7月 20031 8月 2003

学术指纹

探究 'Building a Web Thesaurus from Web Link Structure' 的科研主题。它们共同构成独一无二的学术指纹。

引用此