期刊文献+

基于广义话题理论的话题句识别 被引量:13

Topic Clause Identification Based on Generalized Topic Theory
下载PDF
导出
摘要 汉语标点句句首话题缺失是机器翻译、信息抽取准确率不高的原因之一。该文从广义话题理论出发,根据汉语话题结构的特点,提出标点句的话题句识别研究方案,包括两个阶段性任务:单个标点句的话题句识别和序列标点句的话题句序列构建。识别出标点句的话题句也就找到了标点句句首缺失的话题。该文解决单个标点句的话题句识别任务,主要采用语义泛化和编辑距离两种手段。实验中开放测试的准确率比基线高出12.51个百分点。该结果说明,运用广义话题理论进行单个标点句的话题句识别可产生明显的效果。 Nowadays the Chinese machine translation and information extraction is still far from satisfactory. One important reason is that the topics are often omitted in the head of Chinese Punctuation Clause (abbreviated as PClause). Based on the Generalized Topic Theory, this paper proposes a novel method for topic clause identification from PClause based on the characteristic of topic strcture. The method consists of two tasks in practice: topic clause identification from a single PClause and topic clause construction for a series of PClauses. In the first task,semantic generalization and edit distance are applied in this paper, and the accuracy rate for open test is 12.51% higher than baseline. The result proves the effectiveness of the generalized topic theory in topic clause identification from a single PClause.
作者 蒋玉茹 宋柔
出处 《中文信息学报》 CSCD 北大核心 2012年第5期114-119,128,共7页 Journal of Chinese Information Processing
基金 国家自然科学基金资助项目(60872121,60873013) 北京信息科技大学校基金资助项目(J0725019)
关键词 标点句 广义话题 话题结构 话题句 话题句识别 punctuation clause generalized topic discourse structure topic clause, topic clause identification
  • 相关文献

参考文献2

二级参考文献32

共引文献34

同被引文献139

引证文献13

二级引证文献35

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部