期刊文献+

中文分词和词性标注联合模型综述 被引量:1

The Review on the Joint Model of Chinese Word Segmentation and Part-of-speech Tagging
下载PDF
导出
摘要 中文分词和词性标注任务作为中文自然语言处理的初始步骤,已经得到广泛的研究。由于中文句子缺乏词边界,所以中文词性标注往往采用管道模式完成:首先对句子进行分词,然后使用分词阶段的结果进行词性标注。然而管道模式中,分词阶段的错误会传递到词性标注阶段,从而降低词性标注效果。近些年来,中文词性标注方面的研究集中在联合模型。联合模型同时完成句子的分词和词性标注任务,不但可以改善错误传递的问题,并且可以通过使用词性标注信息提高分词精度。联合模型分为基于字模型、基于词模型及混合模型。本文对联合模型的分类、训练算法及训练过程中的问题进行详细的阐述和讨论。 Chinese word segmentation and part - of - speech (POS) tagging task as an initial step for Chinese natural lan- guage processing, has been widely studied. Due to the lack of Chinese sentences word boundary, the Chinese POS tagging task is often completed with the pipeline approach: firstly, perform Chinese word segmentation, and then use the results of the prior stage to tag the Chinese sentence. However, in the pipeline approach, word segmentation phase errors will be pas- sed to the POS tagging stage, thereby reducing the accuracy of POS tagging. In recent years, the research on Chinese POS tagging focused on the joint model. The joint model perform both word segmentation and POS tagging in a combined single step simultaneously, through which the error propagation can be avoided and the accuracy of word segmentation can be im- proved by utilizing POS information. There are character - based methods, word - based methods, and hybrid methods. In this paper, the three kinds of joint model, the training algorithm and the problems through the processing will be introduced in detail.
出处 《智能计算机与应用》 2014年第3期77-80,共4页 Intelligent Computer and Applications
基金 国家自然科学基金(60975077)
关键词 中文分词 中文词性标注 联合模型 Chinese Word Segmentation Chinese Part- of- speech Tagging Joint Model
  • 相关文献

同被引文献14

引证文献1

二级引证文献44

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部