期刊文献+

自动抽取web数据的树对齐算法

Automatic web data extraction based on tree alignment
下载PDF
导出
摘要 针对从模板生成的网页中自动抽取web数据的问题,提出了一种新的树对齐算法.该算法能够确定输入网页的最大匹配结构.经过一系列的对齐操作之后,多棵树被合并成为一棵记录着合并前多个网页上的统计信息的合并树,树对齐算法可以发现合并树中的重复模式,在最可能内容块上构建包装器,并按照重复模式从网页上抽取数据.实验结果表明,该算法的抽取结果具有较高的准确性和良好的稳定性. This paper proposed a new tree alignment algorithm for determining the optimal matching structure of the input web pages, in order to extract web data automatically. Based on the alignment, the trees were merged into one union tree whose nodes record statistical information obtained from multiple web pages. The algorithm detects repeating patterns on the union tree, and a wrapper built on the most probable content block and the repeating patterns extracts data from web pages. Experimental results showed that the proposed algorithm achieves high extraction accuracy and has steady performance.
出处 《华东师范大学学报(自然科学版)》 CAS CSCD 北大核心 2010年第5期96-102,共7页 Journal of East China Normal University(Natural Science)
关键词 数据抽取 包装器 树对齐 data extraction wrapper tree alignment
  • 相关文献

参考文献9

  • 1刘兵.Web数据挖掘[M].北京:清华大学出版社,2009.
  • 2CHANG C H,KAYED M,GIRGIS M R,et al.A survey of web information extraction systems[J].IEEE Transactions on Knowledge and Data Engineering,2006,18(10):1411-1428.
  • 3徐云风,蒋文蓉.Web页面信息抽取的分析与研究[C]//第十一届中国Java技术及应用交流大会文集.北京:[出版者不详]:2008.
  • 4CRESCENZI V,MECCA G,MERIALDO P.Roadrunner:Towards automatic data extraction from large web.sites[C]// Proc of the 26th Intl Conference on Very Large Database Systems.Rome:[s.n.],2001:109-118.
  • 5ARASU A,HECTOR G M.Extracting structured data from web pages[CJ// Proc of the 2003 ACM SIGMOD Intl Conference on Management of Data.San Diego:[s.n.],2003:337-348.?.
  • 6ZHAI Y,LIU B.Web data extraction based on partial tree alignment[C]// Proc of the 14th Intl World Wide Web Conference(WWW'05).Chiba:[s.n.],2005:76-85.
  • 7ZIGORIS P,EADS D,ZHANG Y.Unsupervised learning of tree alignment models for information extraction[C]// Proc of the 6th IEEE Intl Conference on Data Mining-Workshops.Hong Kong:[s.n.],2006:45-49.
  • 8REIS D C,GOLGHER P B,SILVA A S,et al.Automatic web news extraction using tree edit distance[C]//Proc of the 13th Intl Conference on World Wide Web.New York:[s.n.],2004:502-511.
  • 9韩家炜.数据挖掘:概念与技术[M].北京:机械工业出版社,2007:188-198.

共引文献44

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部