期刊文献+

高维分类型数据加权子空间聚类算法 被引量:2

Algorithm for high-dimensional categorical data weighted subspace clustering
下载PDF
导出
摘要 子空间聚类是高维数据聚类的一种有效手段,子空间聚类的原理就是在最大限度地保留原始数据信息的同时用尽可能小的子空间对数据聚类。在研究了现有的子空间聚类的基础上,引入了一种新的子空间的搜索方式,它结合簇类大小和信息熵计算子空间维的权重,进一步用子空间的特征向量计算簇类的相似度。该算法采用类似层次聚类中凝聚层次聚类的思想进行聚类,克服了单用信息熵或传统相似度的缺点。通过在Zoo、Votes、Soybean三个典型分类型数据集上进行测试发现:与其他算法相比,该算法不仅提高了聚类精度,而且具有很高的稳定性。 Subspace clustering is a kind of effective strategy to high-dimensional data clustering, the principle of subspace clustering is as well as possible keeping original data information, meanwhile as small as possible using subspace to data clustering. Based on the studying of the existing soft subspace clustering, it proposes a new algorithm for subspace searching. The algorithm combines with the size of cluster and information entropy, defines a new subspace dimensional weight distribution mode, and then uses the feature vector of cluster subspace to measure the similarity of two clusters. It uses the idea of agglomerative hierarchical clustering in hierarchical clustering to data clustering, which overcoming the shortcomings of using information entropy or traditional similarity separately. Through the test in the Zoo, Votes, Soybean three typical categorical data set to find out that compared with other algorithms, the proposed algorithm not only can improve the accuracy of clustering, but also has the very high stability.
机构地区 汕头大学工学院
出处 《计算机工程与应用》 CSCD 2014年第23期131-135,202,共6页 Computer Engineering and Applications
基金 国家自然科学基金(No.61170130)
关键词 高维数据 聚类 子空间 信息熵 层次聚类 high-dimensional data clustering subspace information entropy hierarchical clustering
  • 相关文献

参考文献21

  • 1Han Jiawei,Kamber M.数据挖掘概念与技术[M].范明,孟小峰,译.北京:机械工业出版社,2008.
  • 2Iam-On N,Boongoen T,Garrett S,et al.A link-based cluster ensemble approach for categorical data clustering[J].IEEE Knowledge and Data Engineering,2012,24(3):413-425.
  • 3Berkhin P.Survey of clustering data mining techniques[M]//Grouping multidimensional data:recent advances in clustering.Berlin:Springer-Verlag,2006:25-71.
  • 4Macqueen J.Some methods for classification and analysisof multivariate observations[C]//Proceedings of the 5thBerkeley Symposium on Mathematical Statistics and Probability.Berkeley:University of California Press,1967:281-297.
  • 5Kaufman L,Rousseeuw P J.Finding groups in data:anintroduction to cluster analysis[M].Hoboken:John Wiley&Sons Inc,1990:68-72.
  • 6Parsons L,Haque E,Liu H.Subspace clustering for highdimensional data:a review[J].ACM SIGKDD Explorations Newsletter,2004,6(1):90-105.
  • 7Zaki M J,Peters M,Assent I,et al.CLICKS:an effectivealgorithm for mining subspace clusters in categorical datasets[C]//Proceedings of the 11th ACM SIGKDD International Conference on Knowledge Discovery in Data Mining,2005:736-742.
  • 8Tan S,Cheng X,Ghanem M,et al.A novel refinementapproach for text categorization[C]//Proceedings of theACM 14th Conference on Information and KnowledgeManagement,2005:469-476.
  • 9Jing L,Ng M K,Huang J Z.An entropy weighting k-meansalgorithm for subspace clustering of high-dimensionalsparese data[J].IEEE Trans on Knowledge and Data Engineering,2007,19(8):1026-1041.
  • 10Huang J Z,Ng M K,Rong H,et al.Automated variable weighting in k-means type clustering[J].IEEETrans on Pattern Analysis and Machine Intelligence,2005,27(5):657-668.

二级参考文献28

  • 1吴万齐.中国的住房问题及对策[J].城市规划,1988,12(1):13-17. 被引量:2
  • 2阳琳贇,王文渊.聚类融合方法综述[J].计算机应用研究,2005,22(12):8-10. 被引量:28
  • 3杨善林,李永森,胡笑旋,潘若愚.K-MEANS算法中的K值优化问题研究[J].系统工程理论与实践,2006,26(2):97-101. 被引量:189
  • 4EVERITT B S, LANDAU S, LEESE M. Cluster analysis[M]. 4th ed. London: Arnold, 2001.
  • 5JAIN A K, MURTY M N, FLYNN P J. Data clustering: a review [J]. ACM Computing Surveys, 1999,31 ( 3 ) :264-323.
  • 6FRED A L. Finding consistent clusters in data partitions [ C ]//Proc of the 2nd International Workshop on Multiple Classifier Systems. Cambridge: Springer, 2001 : 309-318.
  • 7STREHL A, GHOSH J. Cluster ensembles: a knowledge reuse frame-work for combining multiple partitions [ J ]. Journal of Machine Learning Research, 2003,3(3):583-617.
  • 8HE Zeng-you, XU Xiao-fei, DENG Sheng-chun. A cluster ensemble method for clustering categorical data [ J ]. Information Fusion,2005, 6(2) :143-151.
  • 9FRED A, JAIN A K. Data clustering using evidence accumulation [ C]//Proc of the 16th International Conference on Pattern Recognition. Washington DC : IEEE Computer Society,2002 : 276-280.
  • 10LI Tao-ying, CHEN Yan. Fuzzy clustering ensemble algorithm for partitioning categorical data[ C ]//Proc of the 2nd International Conference on Business Intelligent and Financial Engineering. Washington DC : IEEE Computer Society,2009 : 170-174.

共引文献21

同被引文献14

  • 1Han J W.Kamber M.数据挖掘:概念与技术[M].范明,孟晓峰,译.第3版.北京:机械工业出版社,2012:288-289.
  • 2Bouguessa M,Wang S,Jiang Q.A K-means-based algorithm for projective clustering[C]//Pattern Recognition,2006.ICPR 2006.18th International Conference on.IEEE,2006,1:888-891.
  • 3Ng R T,Han J.CLARANS:A method for clustering objects for spatial data mining[J].Knowledge and Data Engineering,IEEE Transactions on,2002,14(5):1003-1016.
  • 4Ester M,Kriegel H P,Sander J,et al.Density-based spatial clustering of applications with noise(DBSCAN)[C]//Proceedings of 2nd International Conference on Knowledge Discovery and Data Mining(KDD-96).1996,96:226-231.
  • 5Beyer K,Goldsteein J,Ramakrishnan R,et al.When is“nearest neighbors”meaningful?[M]//Database Theory-ICDT’99 springer,Berlin Heidelerg,1999:217-235.
  • 6Agrawal R,Gehrke J,Gunopulos D,et al.Automatic subspace of high dimensional data for data mining application[C]//Proceeding of the 1998 ACM-SIGMOD International Conference on Management of Data,New York,USA,1998:94-105.
  • 7UCI data base[EB/OL].[2012-12-19]http://archive.ics.uci.edu/ml/datasets.html.
  • 8He Z Y,Xu X F,Deng S C.A cluster ensemble method for clustering categorical data[J].Information Fusion,2005,6(2):143-151.
  • 9San O M,Huynh V,Nakamori Y.An alternative extension of the k-means algorithm for clustering categorical data[J].International Journal of Applied Mathematics and Computer Science,2004,14(2):241-247.
  • 10单世民,王新艳,张宪超.高维分类属性的子空间聚类算法[J].小型微型计算机系统,2009,30(10):2016-2021. 被引量:6

引证文献2

二级引证文献6

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部