期刊文献+

相似文本的快速搜索 被引量:1

Faster Algorithm for Searching Similar Text
下载PDF
导出
摘要 相似文本的快速搜索是大规模文本处理需要解决的基本问题。从两方面改进了Udi的相似文本搜索方法,通过Hash把集合映射成ID,从而得到更快的集合比较算法,重新定义了相似关系,能够减少误判,同时对有固定格式的文本也有更好的效果。 Searching similar texts is a fundamental problem for many large scale text processing tasks. Udi's algorithm for searching similar texts is improved in two ways. By mapping each set to an [D, a faster algorithm to compare sets is obtained. And the relation of similar to is redefined in order to both reduce false decision and improve the performance for those texts with fixed format.
出处 《计算机工程》 CAS CSCD 北大核心 2004年第15期22-23,71,共3页 Computer Engineering
基金 国防预研基金资助项目
关键词 大规模文本处理 相似文本搜索 复制检测 Large scale text processing Similar texts searching Copy detection
  • 相关文献

参考文献5

  • 1Heintze N. Scalable Document Fingerprinting. Oakland. California:Proceedings of the Second USENIX Workshop on Electronic Commerce, http://www.cs.cmu.edu/afs/cs/user/nch/www/koa la/main.html,1996
  • 2Rivest. The MD5 Message-Digest Algorithm. http://www.faqs.org/rfcs/rfc 1321 .html, 1992
  • 3Brin S, Davis J, Garcia-Molina H. Copy Detection Mechanisms for Digital Documents. San Francisco,CA: Proc. of the ACM SIGMOD Annual Conference, 1995
  • 4Shivakumar N, Gareia-Molina H. SCAM: A Copy Detection Mechanism for Digital Documents. In: Proceedings of the 2nd International Conference in Theory and Practice of Digital Libraries (DL′95),http://wwwdb. stanford.edu/pub/shivakumar/1995/scam.ps, 1995
  • 5Manber U. Finding Imilar Files in a Large File System. San Francisco,CA: Proceedings of the Winter 1994 USENIX Technical Conference,1994

共引文献1

同被引文献5

引证文献1

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部