期刊文献+

基于主题分析的文本分割技术研究 被引量:6

Research on Text Segmentation Based on Topic Analysis
下载PDF
导出
摘要 本文提出一种新颖的文本分割算法,算法首先将待分割文档划分为若干片段的集合,然后构造全文词汇链分析文中描述的多个子主题,并通过构造片段对子主题的覆盖图将描述相同子主题的相似片段归类.针对段落分割点可能落在片段内部的情况,算法对片段进行二次划分.实验表明:在对文档进行主题分析后,算法能够过滤掉与主题无关的特征对分割结果的干扰;构造的片段对子主题的覆盖图融合了相邻及相间片段的相似性,加大了划分的准确度;对片段进行二次划分使得分割的结果更加合理. A novel topic segmentation algorithm is proposed in this paper.This algorithm first partitions text into some blocks.After that it constructs whole-length lexical chains to analyze multiple subtopics of this text.By constructing graph which describes blocks covering subtopics,the similar blocks which describe same subtopic can be classified.In order to solve the situations that segmentation points drop inside blocks,it segments blocks again.Experiment results demonstrate that by analyzing topic of text,this algorithm can remove interferences,which are aroused by irrelative features,from segmentation results.By constructing graph which describes blocks covering subtopics,it can mix similarities of adjacent and disconnected blocks together,and increases segmentation precision.The second segmentation makes segmentation results more reasonable.
出处 《电子学报》 EI CAS CSCD 北大核心 2009年第2期278-284,共7页 Acta Electronica Sinica
基金 国家自然科学基金重点项目(No.60435020) 国家863高技术研究发展计划项目(No.2006AA01Z197 No.2007AA01Z172)
关键词 主题分析 词汇链 知网 二次划分 topic analysis lexical chain HowNet second segmentation
  • 相关文献

参考文献12

  • 1Kauchak D, Chen F. Feature-based segmentation of narrative documents[ A] .Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics [ C ]. USA: ACL Press, 2005.32 - 39.
  • 2Hearst M A. TextTiling: segmenting text into multi-paragraph subtopic passages[ J]. Computational Linguistics, 1997,23 ( 1 ) : 33 - 64.
  • 3Stokes N, Carthy J, Smeatona F. SeLeCT: a lexical cohesion based news story segmentation system[J]. Journal of AI Communications, 2094,17 ( 1 ) : 3 - 12.
  • 4Chen Qingcai, Wang Xiaolong, Liu Bingquan. Subtopic segmentation of Chinese document: an adapted dotplot approach [A]. Proceedings of 2002 International Conference on Machine Learning and Cybemetics[C]. Beijing: IEEE Press, 2002. 1571 - 1576.
  • 5石晶,戴国忠.基于PLSA模型的文本分割[J].计算机研究与发展,2007,44(2):242-248. 被引量:25
  • 6Morris J, Hirst G. Lexical cohesion computed by thesaural relations as an indicator of the structure of the text[ J]. Computational Linguistics, 1991,17(1) :21 - 48.
  • 7Kok Wee Gan, Ping Wai Wong. Annotating information structures in Chinese texts using HowNet [A]. Proceedings of the Second Workshop on Chinese Language Processing: Held in Conjunction with the 38th Annual Meeting of the Association for Computational Linguistics[ C ]. HK: ACL Press, 2000.85 - 92.
  • 8Chan S W. Extraction of salient textual patterns: synergy between lexical cohesion and contextual coherence [J].IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans,2004,34(2) :205 - 218.
  • 9Gonenc E, Ilyas C. Using lexical chains for keyword extraction [J].Information Processing and Management, 2007, 43 ( 6 ) : 1705- 1714.
  • 10朱靖波,叶娜,罗海涛.基于多元判别分析的文本分割模型[J].软件学报,2007,18(3):555-564. 被引量:15

二级参考文献49

  • 1Lucien Wald. Some terms of reference in data fusion [J]. IEEE Trans Geosci Remote Sensing, 1999,(5):1190-1193.
  • 2Jorge Nunez. Multiresolution-based image fusion with additive wavelet decomposition[J]. IEEE Trans Geosci Remote Sensing, 1999, (5):1204-1211.
  • 3Yocky D A. Image merging and data fusion using discrete two dimensional wavelet transform[J]. Jounal of Opt Soc Am A, 1995, 12(9):1834-1841.
  • 4Lagendijk R L, Biemond J,et al. Identification and restoration of noise blurred images using the expectation-maximization algorithm[J]. IEEE Traus ASSP, 1990,38(7):1150-1191.
  • 5F Maes, A Collignon, D Vandermetden, et al. Multimodality image registration by maximization of mutual information[J]. IEEE Trans Medical Imaging, 1997,16:187-198.
  • 6C Shekhar, V Govindu. R Chellapa. Multisesor image resistration by feature consensus[J]. Pattem Recognition, 1999,32: 39-52.
  • 7Salton G,Singhal A,Buckley C,Mitra M.Automatic text decomposition using text segments and text themes.In:Bernstein M,Carr L,Osterbye K,eds.Proc.of the 7th ACM Conf.on Hypertext.New York:ACM Press,1996.53-65.
  • 8Hearst MA.TextTiling:Segmenting text into multi-paragraph subtopic passages.Computational Linguistics,1997,23(1):33-64.
  • 9Morris J,Hirst G.Lexical cohesion computed by thesauri relations as an indicator of the structure of text.Computational Linguistics,1991,17(1):21-42.
  • 10Kozima H.Text segmentation based on similarity between words.In:Proc.Of the 31st Annual Meeting of the Association for Computational Linguistics.1993.286-288.Http://acl.ldc.upenn.edu/P/P93/P931041.pdf

共引文献33

同被引文献62

引证文献6

二级引证文献57

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部