摘要
介绍了一种集合了规则、串频统计和中文上下文关系分析的现代汉语分词系统.系统对原文进行三次扫描,首先将原文读入内存,利用规则将原文变成若干个串,构成语段十字链表;然后对每个串中的子串在上下文中重复出现的次数进行统计,把根据统计结果分析出的最有可能是词的子串作为临时词;最后利用中文语法的上下文关系并结合词典对原文进行分词处理.系统对未登录词的分词有很好的效果.
A modern Chinese character segmentation system based on rule, statistics and context analysis is described. The system scans the article 3 times. At the first time,it reads the article into memory and then divides it into phases and makes it into intercrossing link by using rules. At the second time,it counts the times that the strings appear. At the last time,with the help of large amount of statistical data and the grammar of the Chinese,it segments Chinese character. It is shown that the system has good performance on the unregistered words.
出处
《内蒙古师范大学学报(自然科学汉文版)》
CAS
2008年第1期71-74,共4页
Journal of Inner Mongolia Normal University(Natural Science Edition)
基金
四川省教育厅重点科研基金资助项目(2003A105)
云南省计算机技术应用重点实验室开放基金资助项目
关键词
中文分词
未登录词
现代汉语自动分词系统
Chinese segmentation
unknown word
modern Chinese character segmentation system