摘要
相似重复记录识别是数据清理中的一个关键问题。文章针对常用的多趟邻接排序法提出了两点改进:一是在多趟排序识别过程中直接合并有重叠的相似记录集,取消了最后计算传递闭包的环节;二是利用关键字按字典序排序的特性,在求编辑距离之前先过滤前面的公共子串,减少了相似记录比较的开销。文章最后给出了改进算法与原算法的对比试验结果。
Detecting approximately duplicate database records is an important task in data cleaning. A new duplicate detection methods was proposed in this paper which improved the familiar MPN method in two ways. Firstly, the step of computing transitive closure was canceled by directly uniting the overlapping similar record sets. Secondly, the cost of records comparing was reduced by filtering the former common substring before computing the edit distance of two keywords. The experimental results were given out between the imoroved algorithm and MPN,
出处
《微计算机信息》
北大核心
2005年第08X期147-149,3,共4页
Control & Automation
关键词
数据清理
相似重复记录
字符串匹配
MPN
传递闭包
Data cleaning, Approximately duplicate databaserecords, String matching, MPN, Transitive closure.