基于不完备标签数据的半监督聚类算法

Semi-Supervised Clustering Algorithm Based on Incomplete Labeled Data

下载PDF

导出

摘要针对seeded-K-means和constrained-K-means算法要求标签数据类别完备的限制,本文提出了基于不完备标签数据的半监督K-means聚类算法,重点讨论了未标签类别初始聚类中心的选取问题。首先给出了未标签类别聚类中心最优候选集的定义,然后提出了一种新的未标签类别初始聚类中心选取方法,即采用K-means算法从最优候选集中选取初始聚类中心,最后给出了基于新方法的半监督聚类算法的完整描述,并通过实验测试对新算法的有效性进行了验证。实验结果表明本文所提算法在执行速度和聚类效果上都优于现有算法。 For the seeded-K-means and constrained-K-means algorithm limitations that complete category information in labeled data is required, this paper put forword an semi-supervised K-means clustering algorithm based on incomplete labeled data, focused on selection of the initial cluster center of unlabeled category. We gave a definition of the Best Candidate Set of cluster center of unlabeled category, proposed a new method that selecting initial cluster center of unlabeled category from the Best Candidate Set using K-means. Finally, a complete description of semi-supervised clustering algorithm based on the new method is given, the validity of the new algorithm is verified by experiment. Experimental results show that the proposed algorithm is superior to existing algorithms not only in clustering effect and in execution speed.

作者袁利永

机构地区浙江师范大学数理与信息工程学院

出处《计算机系统应用》 2011年第2期182-185,共4页 Computer Systems & Applications

关键词半监督聚类:K.means 不完备先验知识初始聚类中心标签数据 semi-supervised clustering K-means incomplete prior knowledge initial cluster center labeled data

分类号 TP311.13 [自动化与计算机技术—计算机软件与理论]

引文网络
相关文献

参考文献10

1Pedrycz W, Vukovich G: Fuzzy. clustering with supervision.. Pattern Recognition, 2004,37(7): 1339 - 1349.
2Wagstaff K, Cardie C, Rogers Clustering with Background S, et al. Constrained K-Means Knowledge. In: Brodley CE, Danyluk AP, eds. Proc. of the 18th lnt'l Conf. on Machine Learning. Williamstown: Morgan Kaufmann Publishers, 2001. 577-584.
3Basu S, Banerjee A, Mooney RJ. Semi-Supervised clustering by seeding. In: Claude S, Achim GH, eds. Proc. of the 19th Int'l Conf. on Machine Learning(ICML2002). San Fransisco: Morgan Kaufxnann Publishers, 2002. 19-26.
4李志圣,孙越恒,何丕廉,侯越先.基于k-means和半监督机制的单类中心学习算法[J].计算机应用,2008,28(10):2513-2516. 被引量：4
5高滢,刘大有,齐红,刘赫.一种半监督K均值多关系数据聚类算法[J].软件学报,2008,19(11):2814-2821. 被引量：22
6Kulis B, Basu S, Dhillon I, et al. Semi-Supervised Graph Clustering: A Kernel Approach. Machine Leaming, 2009, 1 (74): 1 - 22.
7Jain AK, Dubes RC. Algorithms for clustering data. Englewood Cliffs, N J: Prentice Hall, 1988.
8Han J, Kamber M. Data mining: Concepts and techniques. SanFrancisco: Morgan Kaufmann, 2001.
9MacQueen J. Some methods for classification and analysis of multivariate observations. Proc. of the 5th Berkeley Symposium on Mathematical Statistics and Probability. Berkeley, CA, 1967,1:281 -297.
10高云天,王学辉,郭涛.基于不完整信息的半监督聚类算法[J].北华大学学报（自然科学版）,2009,10(5):457-463. 被引量：2

二级参考文献28

1Dzeroski S. Multi-Relational data mining: An introduction. ACM SIGKDD Explorations Newsletter, 2003,5(1):1-16.
2Dzeroski S, Lavrac N. Relational Data Mining. Berlin: Springer-Verlag, 2001. 339-364.
3Domingos P. Prospects and challenges for multi-relational data mining. ACM SIGKDD Explorations Newsletter, 2003,5(1):80-83.
4Bouchachia A. Learning with partly labeled data. Neural Computing and Applications, 2007,16(3):267-293.
5Zhu XJ. Semi-Supervised learning literature survey. Technical Report, Computer Sciences TR 1530, University of Wisconsin- Madison, 2007. 1-42.
6Chapelle O, Seholkopf B, Zien A. Semi-Supervised Learning. Cambridge: MIT Press, 2006. 3-14.
7Long B, Zhang F, Wu XY, Yu PS. Spectral clustering for multi-type relational data. In: Cohen WW, Moore A, eds. Proc. of the 23rd Int'l Conf. on Machine Learning. New York: ACM Press, 2006. 585-592.
8Marques de Sa JP, Wrote; Wu YF, Trans. Pattern Recognition Concepts, Methods and Applications. 2nd ed., Beijing: Tsinghua University Press, 2002.51-74 (in Chinese).
9http://archive.ics.uci.edu/ml/datasets.html
10Yin XX, Han JW, Yu PS. CrossClus: User-Guided multi-relational clustering. Data Mining Knowledge Discovery, 2007,15(3): 321-348.

共引文献23

1孙雪,李昆仑,胡夕坤,赵瑞.基于半监督K-means的K值全局寻优算法[J].北京交通大学学报,2009,33(6):106-109. 被引量：11
2孙晓鹏,张琪,魏小鹏.半监督的三维网格模型层次分割[J].计算机辅助设计与图形学学报,2010,22(4):592-598. 被引量：5
3李小展.基于半监督的K-means聚类改进算法[J].东莞理工学院学报,2011,18(1):29-32. 被引量：1
4杨南海,黄明明,赫然,王秀坤.基于最大相关熵准则的鲁棒半监督学习算法[J].软件学报,2012,23(2):279-288. 被引量：8
5芦世丹,崔荣一.基于主动学习策略的半监督聚类算法研究[J].计算机应用研究,2013,30(6):1718-1720. 被引量：1
6梅松青.基于自适应图的半监督学习方法[J].计算机系统应用,2014,23(2):173-177. 被引量：2
7于重重,吴子珺,谭励,涂序彦,杨扬,王璐.多元时序模糊聚类分段挖掘算法[J].北京科技大学学报,2014,36(2):260-265. 被引量：3
8文翰,肖南峰.基于强类别特征近邻传播的半监督文本聚类[J].模式识别与人工智能,2014,27(7):646-654. 被引量：10
9黄少滨,程媛,万庆生,刘国峰,申林山.一种基于IDEF1x模型的层次多关系聚类算法[J].自动化学报,2014,40(8):1740-1753. 被引量：1
10谢梦燕,黄旭,赵青,王俊辉.一种不规则形状聚类算法[J].西安文理学院学报（自然科学版）,2015,18(3):5-8.

1袁利永,王基一.一种改进的半监督K-Means聚类算法[J].计算机工程与科学,2011,33(6):138-143. 被引量：13
2徐晓丹.基于半监督学习的中文多文档子主题划分[J].浙江师范大学学报（自然科学版）,2011,34(3):302-305. 被引量：1
3周萍,秦永彬,黄瑞章.结合seeds集和LDA的半监督文本聚类算法[J].计算机工程与设计,2014,35(6):1994-1998. 被引量：1
4唐明亮,沈晓冬.Modeling and Estimation of the Kinetics of Seeded Solvent-mediated Phase Transformation in a Batch Crystallizer[J].Journal of Wuhan University of Technology(Materials Science),2011,26(5):872-878. 被引量：1
5杨威,周林,龙世同,彭文翠,王谨,詹明生.Time-division-multiplexing laser seeded amplification in a tapered amplifier[J].Chinese Optics Letters,2015,13(1):49-52.
6马红梅,陈丽清,袁春华.Cascade correlation-enhanced Raman scattering in atomic vapors[J].Chinese Physics B,2016,25(12):271-276.
7单体锋,逄少军,高素芹.Periodic exposure to ambient solar irradiance benefits the growth of juvenile seedlings of Hizikia fusiformis[J].Chinese Journal of Oceanology and Limnology,2011,29(5):1009-1014.
8徐志蓉,姜发纲,曾艳彩,Hamed TM Alkhodari,陈飞.Culture of Rat Retinal Ganglion Cells[J].Journal of Huazhong University of Science and Technology(Medical Sciences),2011,31(3):400-403.

计算机系统应用

2011年第2期

浏览历史

内容加载中请稍等...

基于不完备标签数据的半监督聚类算法

参考文献10

二级参考文献28

共引文献23

相关作者

相关机构

相关主题

浏览历史