[KSCI] Korea Science Citation Index Service

Refining Initial Seeds using Max Average Distance for K-Means Clustering

Lee, Shin-Won (중원대학교 IT공학부)
Lee, Won-Hee (전북대학교 대학원 컴퓨터공학과)

Publication Information

Journal of Internet Computing and Services / v.12, no.2, 2011 , pp. 103-111 More about this Journal

Abstract

Clustering methods is divided into hierarchical clustering, partitioning clustering, and more. If the amount of documents is huge, it takes too much time to cluster them in hierarchical clustering. In this paper we deal with K-Means algorithm that is one of partitioning clustering and is adequate to cluster so many documents rapidly and easily. We propose the new method of selecting initial seeds in K-Means algorithm. In this method, the initial seeds have been selected that are positioned as far away from each other as possible.

Keywords

clustering; K-Means; initial seed;

Citations & Related Records

Reference

1	이신원 "정보검색을 위한 개선된 K-Means 알고리즘을 이용한 계층적 클러스터링에 관한 연구", 박사학위 논문, 전북대학교, 2005.
2	Paul Bunn, and Rafail Ostrovsky, "Secure Two-Party k-Means Clustering", Proceedings of the 14th ACM conference on Computer and communications security, Alexandria, Virginia, USA, pp.486-497, 2007.
3	Rafail Ostrovsky, Yuval Rabani, Leonard J. Schulman and Chaitanya Swamy, "The Effectiveness of Lloyd-Type Methods for then k-Means Problem", Proceedings of the 47th Annual IEEE Symposium on Foundaions of Computer Science, pp.165-176, 2006.
4	Nachiketa Sahoo, Jamie Callan, Ramayya Krishnan, George Duncan, and Rema Padman, "Incremental hierarchical clustering of text documents", Proceedings of the 15th ACM international conference on Information and knowledge management, pp.357-366, 2006.
5	Yu Yonghong, and Bai Wenyang, "Text clustering based on term weights automatic partition", Computer and Automation Engineering (ICCAE), 2010 The 2nd International Conference, pp.373-377, 2010.
6	Giordano Adami, Paolo Avesani, and Diego Sona, "Clustering documents in a web directory", Proceedings of the 5th ACM international workshop on Web information and data management, pp.66-73, 2003.
7	Christopher D. Manning, Prabhakar Raghavan and Hinrich Schütze, "Introduction to Information Retrieval", Cambridge University Press, pp.331-338, 2008.
8	Jain, A. K. and Dubes, R. C., "Algorithms for Clustering Data". Prentice-Hall advanced reference series. Prentice-Hall, Inc., Upper Saddle River, NJ. 1988.
9	McQueen, J. "Some methods for classification and analysis of multivariate observations", In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, pp.281-297, 1967.
10	S. P. Lloyd, "Least squares quantization in PCM", Special issue on quantization, IEEE Trans. Inform. Theory, 28, pp.129-137, 1982. DOI
11	D.A.Meedeniya, and A.S.Perera, "Evaluation of Partition-Based Text Clustering Techniques to Categorize Indic Language Documents", IEEE International Advance Computing Conference (IACC 2009), pp.1497-1500, 2009.

1	Comparison of Initial Seeds Methods for K-Means Clustering / [Lee, Shinwon;] / Journal of Internet Computing and Services
2	A Road Extraction Algorithm using Mean-Shift Segmentation and Connected-Component / [Lee, Tae-Hee;Hwang, Bo-Hyun;Yun, Jong-Ho;Park, Byoung-Soo;Choi, Myung-Ryul;] / Journal of Digital Convergence

KSCI

Refining Initial Seeds using Max Average Distance for K-Means Clustering K-Means 클러스터링 성능 향상을 위한 최대평균거리 기반 초기값 설정

Refining Initial Seeds using Max Average Distance for K-Means Clustering