期刊文献+

一种基于同义词扩展的不平衡文本分类方法 被引量:1

An Imbalanced Text Classification Method Based on Synonyms Expansion
下载PDF
导出
摘要 针对传统文本分类方法的性能,尤其是其中少数类的分类性能会随着文本不平衡程度的加重而迅速恶化的现象,提出了一种基于同义词扩展的不平衡文本分类改进方法。该方法通过建立同义词词典、确定扩展规则和调整"特征保持因子"等几个步骤,实现了少数类中的特征项的丰富和补偿,同时对扩展带来的原文档特征变化予以了补偿。实验结果表明,该方法可以从很大程度上改善少数类的分类性能,并且随着少数类中文本数量的减少,性能的提升会越发显著。与此同时,分类器的总体分类性能也得到了一定程度的提升。 The performance of traditional text categorization methods, especially the categorization performance for minority classes, often deteriorates rapidly for imbalanced text. A new method based on synonyms expansion is introduced in this paper in order to deal with im- balanced text classification. With the steps of the establishment of synonym-dictionary, the determination of expansion rules and the modi- fication of the "Feature-Maintaining Factor", feature items of minority classes are enriched. At the same time, the changes brought by the expansion are compensated. The experimental results show the categorization performance for minority classes is improved to a high de- gree. Moreover, with the decrease of the quantity of text in minority classes, the performance improves significantly. The overall perform- ance is improved to some degree at the same time.
出处 《情报杂志》 CSSCI 北大核心 2013年第9期204-206,F0003,共4页 Journal of Intelligence
基金 国家自然科学基金项目"基于行为分析的网络流量检测技术研究"(编号:60972077)的资助
关键词 文本分类 不平衡数据集 同义词词典 词频保持 text classification imbalanced dataset synonym-dictionary term-frequency maintaining
  • 相关文献

参考文献9

  • 1Wasikowski M, Xue-wen Chen. Combating the Small SampleClass Imbalance Problem Using Feature Selection[ J]. Knowl-edge and Date Engineering, 2010,22(10) :1388-1400.
  • 2廖一星,潘雪增.面向不平衡文本的特征选择方法[J].电子科技大学学报,2012,41(4):592-595. 被引量:5
  • 3Mladenic D, Grobdnik M. Feature Selection for Unbalanced CassDistribution and Naive Bayes[ C]. Proceedings of 16th Interna-tional Conference on Machine Learning. San Francisco, 1999:258-267.
  • 4Zhengyu Lu, Yongmin Lin, Shuang Zhao, etl al. Study on Fea-ture Selection and Weighting Based on Synonym Merge in TextCategorization[ C]. Proceedings of 2nd International Conferenceon Feature Networks, 2010:105-109.
  • 5刘群,张华平,俞鸿魁,程学旗.基于层叠隐马模型的汉语词法分析[J].计算机研究与发展,2004,41(8):1421-1429. 被引量:197
  • 6Yiming Yang, Jan O Pedersen. A Comparative Study on FeatureSelection in Text Categorization [ C]. Proceedings of 14th Inter-national Conference on Machine Learning. Nashville, 1997:412-420.
  • 7Yan Xu. A Comparative Study on Feature Selection in ChineseSpam Filtering[C]. Proceedings of 6th International Conferencen Application of Information and Communication Technologies.Tbilisi, 2012:1-6.
  • 8Chua S,Kulathuramaiyer N. Semantic Feature Selection Usingwordnet [ C]. Proceedings of IEEE/W1C/ACM InternationalConference on Wed Intelligence. WI, 2004:166-172.
  • 9任纪生,王作英.基于特征有序对量化表示的文本分类方法[J].清华大学学报(自然科学版),2006,46(4):527-529. 被引量:4

二级参考文献45

  • 1徐燕,李锦涛,王斌,孙春明,张森.不均衡数据集上文本分类的特征选择研究[J].计算机研究与发展,2007,44(z2):58-62. 被引量:20
  • 2H Y Tan. Chinese place automatic recognition research. In: C N Huang, Z D Dong, eds. Proc of Computational Language.Beijing: Tsinghua University Press, 1999
  • 3Zhang Huaping, Liu Qun, Zhang Hao, et al. Automatic recognition of Chinese unknown words recognition. First SIGHAN Workshop Attached with the 19th COLING, Taipei, 2002
  • 4S R Ye, T S Chua, J M Liu. An agent-based approach to Chinese named entity recognition. The 19th Int'l Conf on Computational Linguistics, Taipei, 2002
  • 5J Sun, J F Gao, L Zhang, et al. Chinese named entity identification using class-based language model. The 19th Int'l Conf on Computational Linguistics, Taipei, 2002
  • 6Lawrence R Rabiner. A tutorial on hidden Markov models and selected applications in speech recognition. Proc of IEEE, 1989,77(2): 257~286
  • 7Shai Fine, Yoram Singer, Naftali Tishby. The hierarchical hidden Markov model: Analysis and applications. Machine Learning,1998, 32(1): 41~62
  • 8Richard Sproat, Thomas Emerson. The first international Chinese word segmentation bakeoff. The First SIGHAN Workshop Attached with the ACL2003, Sapporo, Japan, 2003. 133~143
  • 9J Hockenmaier, C Brew. Error-driven learning of Chinese word segmentation. In: J Guo, K T Lua, J Xu, eds. The 12th Pacific Conf on Language and Information, Singapore, 1998
  • 10Andi Wu, Zixin Jiang. Word segmentation in sentence analysis.1998 Int'l Conf on Chinese Information Processing, Beijing, 1998

共引文献201

同被引文献3

引证文献1

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部