期刊文献+

面向中文网络百科的属性和属性值抽取 被引量:12

Attribute and Attribute Value Extracted from Chinese Online Encyclopedia
下载PDF
导出
摘要 针对面向中文网络百科条目文章的属性和属性值抽取,提出一种无监督方法。此方法将属性值看做命名实体,利用频繁模式挖掘和关联分析,从文本中抽取类别属性;采用自扩展方法为属性建立触发词表;基于属性触发词和属性值实体标注挖掘属性值抽取模式,利用层次聚类算法获取高质量的模式。在互动百科中采集的数据集上进行实验,结果表明所提方法行之有效。 An unsupervised approach is proposed to extract attribute and attribute value from Chinese online encyclopedia entry articles. Attribute values are viewed as named entities and class attributes are extracted based on frequent patterns mining and association analysis. A bootstrapping method is used to find attribute trigger words for each attribute. Attribute value extraction patterns are generated automatically from sentences which contain attribute trigger words and named entity tags of attribute value. Hierarchy clustering algorithm is applied to obtain reliable patterns. Experimental dataset are collected from HudongBaike. The experiment results show that the method is feasible and effective.
出处 《北京大学学报(自然科学版)》 EI CAS CSCD 北大核心 2014年第1期41-47,共7页 Acta Scientiarum Naturalium Universitatis Pekinensis
基金 国家自然科学基金(61170111 61202043 61262058) 中国科学院自动化研究所复杂系统管理与控制重点实验室开放课题(20110102) 中央高校基本科研业务费专项基金(SWJTU11ZT08)资助
关键词 知识获取 属性抽取 非结构化文本 模式挖掘 knowledge acquisition attribute extraction unstructured text pattern mining
  • 相关文献

参考文献29

  • 1Suchanek F,Kasneci G,Weikum G. Yago:a core of semantic knowledge unifying WordNet and Wikipedia // Proc of WWW 2007[J].New York:ACM,2007.697-706.
  • 2Auer S,Bizer C,Lehmann G. DBpedia:a nucleus for a Web of open data[A].{H}Berlin:Springer-Verlag,2007.722-735.
  • 3Wu Fei,Weld D. Autonomously Semantifying Wikipedia[A].New York:ACM,2007.41-50.
  • 4Wu Fei,Weld D. Automatically refining the Wikipedia Infobox Ontology[A].New York:ACM,2008.635-644.
  • 5赵军,刘康,周光有,蔡黎.开放式文本信息抽取[J].中文信息学报,2011,25(6):98-110. 被引量:61
  • 6Tokunaga K,Kazama J,Torisawa K. Automatic discovery of attribute words from web documents //Proc of IJCNLP 2005[J].{H}Berlin:Springer-Verlag,2005.106-118.
  • 7Pa(s)ca M. Organizing and searching the world wide web offacts-step two:Harnessing the wisdom of the crowds[A].New York:ACM,2007.101-110.
  • 8Pa(s)ca M,Durme B. Weakly-supervised acquisition of open-domain classes and class attributes from web documents and query logs[A].Stroudsburg:ACL,2008.19-27.
  • 9Kopliku A,Sauvagnat K,Boughanem M. Retrieving attributes using web tables[A].New York:ACM,2011.13-17.
  • 10Sanchez D. A methodology to learn ontological attributes from the web[J].{H}Data & Knowledge Engineering,2010,(69):573-597.

二级参考文献141

共引文献335

同被引文献185

引证文献12

二级引证文献102

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部