期刊文献+

基于向量空间模型的古汉语词义自动消歧研究 被引量:6

Automatic Word Sense Disambiguation of Ancient Chinese Based on Vector Space Model
原文传递
导出
摘要 借鉴现代汉语词义消歧的研究成果,提出一种改进的向量空间模型词义消歧方法,即在古汉语义项词语知识库的支持下,将待消歧多义词上下文与多义词的义项映射到向量空间模型中,完成语义消歧任务。以中国农业古籍全文数据库为统计语料,对10个典型古汉语多义词,共29个义项、1 836条待消歧上下文进行义项标注的实验,消歧平均正确率达到79.5%。 How to annotate the meaning of words is an important research work on collation of Chinese ancient books. The manual interpretation is time -consuming and laborious. According to the word sense disambiguation of modern Chinese, an improved unsupervised disambiguation method of ancient Chinese is proposed based on the vector space model. In order to disambiguate the word sense, the knowledge repository of ancient Chinese polysemous words is build, and the contexts and the meanings of the polysemous words are mapped into the vector space model. This paper takes the full - text database of Chinese agricultural ancient books for statistics corpus, and conducts the experiment using 10 typical polysemous words of ancient Chinese which include 29 senses and 1836 contexts. The result shows that the average disambiguation accuracy achieves 79.5%.
出处 《图书情报工作》 CSSCI 北大核心 2013年第2期114-118,共5页 Library and Information Service
基金 国家社会科学基金项目"古籍整理与开发智能化技术研究"(项目编号:08ATQ002) 高等学校博士学科点专项科研基金资助课题"古农书资料自动编纂及注释系统的设计与构建"(项目编号:20090097110033)研究成果之一
关键词 向量空间模型 词义消歧 古汉语 vector space model semantic disambiguation ancient Chinese
  • 相关文献

参考文献16

  • 1百度百科.古书注解[EB/OL].[2012-05-23].http ://baike. baidu. com/view/793424. htm#3.
  • 2卢志茂,刘挺,李生.统计词义消歧的研究进展[J].电子学报,2006,34(2):333-343. 被引量:27
  • 3Lesk M. Automatic Sense Disambiguation Using Machine Readable Dictionaries: how to tell a pine cone from an ice cream cone[ C ]// Proceedings of the 5th International Conference on Systems Documentation. Toronto Canada: ACM, 1986 : 24 - 26.
  • 4Manning C D, Schutze H. Foundations of statistical natural language processing [ M ]. Cambridge : The MIT Press, 1999 : 229 - 260.
  • 5Yarowsky D. Word-sense disambiguation using statistical models of Roger' s categories trained on large corpora[ EB/OL]. [ 2012 -05 -23 ]. http://www, informatik, uni -trier. de/~ ley/db/conf/coling/coling1992. html.
  • 6Ng H T and Lee H B. Integrating multiple knowledge sources to disambiguate word sense: An example based approach [ EB/OL]. [2012 -05 -23 ]. http://citeseerx. ist. psu. edu/showciting?cid = 4549.
  • 7张仰森,郭江.四种统计词义消歧模型的分析与比较[J].北京信息科技大学学报(自然科学版),2011,26(2):13-18. 被引量:7
  • 8李永亮,黄曙光,鲍蕾.一种基于PageRank算法和知网的词义消歧方法[J].计算机应用与软件,2011,28(5):213-215. 被引量:4
  • 9Lin Shoude, Karin V. A semantics-Enhanced language model for unsupervised word sense disambiguation [ C ]//Proceedings of the 9th International Conference on Computational Linguistics and Intelligent Text Proceeding. Haifa: Springer,2008:287 -298.
  • 10李娟子.汉语词义消歧方法研究[D].北京:清华大学,1999.

二级参考文献107

共引文献77

同被引文献148

引证文献6

二级引证文献40

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部