期刊文献+

基于语义的互联网药品信息抽取算法 被引量:7

Web Medicine Information Extraction Algorithm Based on Semantics
下载PDF
导出
摘要 针对现有互联网信息抽取技术存在准确率不高、覆盖率低、人工干预多等诸多缺陷,提出了一种新的互联网药品信息抽取算法,通过引入语义技术构建三维语义词典,屏蔽不同药品信息网页在内容和结构上的异构性,同时利用所需抽取的目标药品属性信息具有一定聚集度的特征,基于信息熵的理论设计出对目标信息智能定位和抽取的方法。实验证明该算法既能降低人工干预,又具备较高的准确率和召回率。应用该算法能实时自动全面准确地获取互联网药品信息,为政府药监部门提供丰富的监管依据,对规范医药电子商务市场,保证人们的用药安全具有重要的现实意义。 This article addresses defects of current Web information extraction technology such as low accuracy, low coverage, and manual intervention required, proposes a novel extraction algorithm of web medicine information. The algorithm sets up a three-dimentional semantic dictionary by introduction of the semantics technology, masks the isomerisms of the web page contents and structures, and at the same time, taking advantage of the fact that the attributes of the target medicine tend to have a character of aggregation, designs a way of intellectually locating and extracting the target information based on the theory of information entropy. Through related experiments proves that the algorithm is able to reduce the requirement of manual intervention of the information extraction, and has a high accuracy and recall rate. The application of this algorithm can automatically, comprehensively, and accurately obtain Internet medicine information in real time, offers abundant basis of supervision for the medicine supervision department, and therefore has a significant practical meaning of normalizing medical e-business and ensuring secure medication.
出处 《计算机系统应用》 2011年第1期41-47,共7页 Computer Systems & Applications
基金 国家科技支撑项目(2006BAH02A05-06) 国家自然科学基金(60903078 60973025)
关键词 WEB信息抽取 语义词典 DOM 信息熵 XPATH 医药电子商务 Web information extraction semantic dictionary DOM information entropy XPath medical E-business
  • 相关文献

参考文献10

  • 1吴晓彦.基于结构语义熵的互联网商品信息抽取技术研究[D]复旦大学,复旦大学2009.
  • 2Selkow S.The Tree-to-Tree Editing Problem. Journal of the Information Processing Letters . 1977
  • 3Ion Muslea,Steve Minton,Craig Knoblock.A hierarchical approach to wrapper induction. Proceedings of the Third International Conference on Autonomous Agents . 1999
  • 4Valter Crescenzi,Giansalvatore Mecca,Paolo Merialdo.RoadRunner:Towards Automatic Data Extraction from Large Web Sites. Proceedings of the 26th International Conference on Very Large Database Systems . 2001
  • 5D Freitag.Information extraction from HTML: application of a general machine learning approach. Proceedings of the Fifteenth National Conference on Artificial Intelligence . 1998
  • 6Soderland,Stephen.Learning information extraction rules for semi-structured and free text. Machine Learning . 1999
  • 7Chia-Hui Chang,Shao-Chen Lui.IEPAD: information extraction based on pattern discover. Proceedings of the 10th International Conference on the World Wide Web . 2001
  • 8N Kushmerick,DS Weld,RB Doorenbos.Wrapper Induction for Information Extraction. Proceedings of the Fifteenth International Joint Conference on Artificial Intelligence(IJCAI297) . 1997
  • 9Arocena G,Mendelzon A.WebOQL: Restructuring Documents, Databases and Webs. Proceedings of the 14th IEEE International Conference on Data Engineering (ICDE) . 1998
  • 10SAHUGUET A,AZAVANT F.Building intelligent web applications using lightweight wrappers. Data Mining and Knowledge Discovery . 2001

同被引文献65

  • 1吴平博,陈群秀,马亮.基于时空分析的线索性事件的抽取与集成系统研究[J].中文信息学报,2006,20(1):21-28. 被引量:21
  • 2陈再良,徐德智,陈学工,沈海澜.基于链式结构XML文档的生成方法[J].计算机工程,2006,32(20):59-61. 被引量:5
  • 3中华人民共和国突发事件应对法[J].中华人民共和国国务院公报,2007(30):16-23. 被引量:9
  • 4Guidelines for Robot Writers[EB/OL].http://info.webcrawler. com/mak/projects/robots/robots.html(Accessed Jul. 25,2006).
  • 5丁宝琼.网络文本信息采集分析关键技术研究与实现[D].解放军信息工程大学,2010.
  • 6于毅,毛明.一种新的脆弱性漏洞扫描器[J].信息安全与通信保密,2007,29(12):89-90. 被引量:1
  • 7Zhiwei F., 2002, Evolution and Present Situation of Corpus Research In China, Journal of Chinese Lan- guage and Computing, 12(1) .43-62.
  • 8李素芳.《“知之于困学,好之于交流,乐之于应用”—专访梁茂成教授,李文中教授和许家金博士》,《中国英语教育》2010年第1期.
  • 9Zhan Weidong, Chang Baobao, Duan Huiming, Zhang Huarui. 2006, "Recent Developments in Chinese Corpus Re- search", The 13'h NIJL International Symposium, Language Corpora. Their Compliation and Application. Tokyo, Ja- pan. 3.6-7. http .//ccl. pku. edu. cn/doubtfire/papers/2006_Corpora_NIJL Workshop. pdf, 2014 年7 月 11日.
  • 10刘成飞.《汉语中介语语料库中汉字偏误处理的比较研究》,http.//www.doe88.com/p-0116174114179.html,2015年06月11日.

引证文献7

二级引证文献45

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部