跳到主要导航 跳到搜索 跳到主要内容

An n-gram-based approach for detecting approximately duplicate database records

  • Zengping Tian*
  • , Hongjun Lu
  • , Wenyun Ji
  • , Aoying Zhou
  • , Zhong Tian*
  • *此作品的通讯作者
  • Fudan University
  • Hong Kong University of Science and Technology
  • IBM

科研成果: 期刊稿件文章同行评审

摘要

Detecting and eliminating duplicate records is one of the major tasks for improving data quality. The task, however, is not as trivial as it seems since various errors, such as character insertion, deletion, transposition, substitution, and word switching, are often present in real-world databases. This paper presents an n-grambased approach for detecting duplicate records in large databases. Using the approach, records are first mapped to numbers based on the n-grams of their field values. The obtained numbers are then clustered, and records within a cluster are taken as potential duplicate records. Finally, record comparisons are performed within clusters to identify true duplicate records. The unique feature of this method is that it does not require preprocessing to correct syntactic or typographical errors in the source data in order to achieve high accuracy. Moreover, sorting the source data file is unnecessary. Only a fixed number of database scans is required. Therefore, compared with previous methods, the algorithm is more time efficient.

源语言英语
页(从-至)325-331
页数7
期刊International Journal on Digital Libraries
3
4
DOI
出版状态已出版 - 2000
已对外发布

指纹

探究 'An n-gram-based approach for detecting approximately duplicate database records' 的科研主题。它们共同构成独一无二的指纹。

引用此