期刊文献+

一种有效的XML数据清洗方法 被引量:1

Efficient Cleaning Approach for XML Data
下载PDF
导出
摘要 研究XML格式的重复数据元素的特点,提出对于特定应用领域,在具体的上下文环境中主动学习XML重复元素的识别规则。通过结构转换,将结构不尽相同的XML数据映射成结构一致的数据,并通过学习不同层次数据元素间的依赖关系权重来获得匹配规则。根据学习得到的转换和匹配规则,采用哈希过滤的方法来提高检测重复XML元素的效率。该方法能够有效地解决XML重复检测面临的结构多样性的问题,理论分析和实验表明,该方法有较高的精度和效率。 By studying characteristics of duplicate XML data, this paper proposes an active machine learning method for a specific application, which is applied to glean transformation rules and matching rules, and accurately identify duplicate XML elements. Transfomation rules are used to eliminate the structural diversities among elements and matching rules are used to identify the relationships between parent and child nodes. In turn, during the detection phase an efficient hash filter algorithm is proposed to reduce computational complexity. Theory and experiment shows that the method can solve this problem efficiently and effectively.
出处 《计算机工程》 CAS CSCD 北大核心 2008年第15期47-50,共4页 Computer Engineering
基金 江苏省"十五"高科技计划基金资助项目(BG2001013)
关键词 主动学习 匹配规则 哈希 active learning matching rules hash
  • 相关文献

参考文献5

  • 1Weis M, Naumann F. Detecting Duplicate Objects in XML Documents[C]//Proceedings of the 2004 International Workshop on Information Quality in Information Systems. Paris, France: [s. n.], 2004: 10-19.
  • 2Tejada S, Knoblock C A, Minton S. Learning Object Identification Rules for Information Integration[D]. CaliFornia, USA: University of Southern California, 2002.
  • 3Breiman I. Bagging Predictors Machine Learning[J]. 1996, 24(2): 123-140.
  • 4Sung S Y, Li Zhao, Sun Peng. A Fast Filtering Scheme for Large Database Cleansing[C]//Proceedings of the 11th International Conference on Information and Knowledge Management. Virginia, USA: [s. n.], 2002: 76-83.
  • 5Ukkonen E. Approximate String-matching with Q-grams and Maximal Matches[J]. Theore-tical Computer Science, 1992, 92(1): 191-212.

同被引文献7

引证文献1

二级引证文献3

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部