[KSCI] Korea Science Citation Index Service

http://dx.doi.org/10.6109/jkiice.2014.18.3.625

Document Clustering Technique by K-means Algorithm and PCA

Kim, Woosaeng (Department of Computer Software, Kwangwoon University)
Kim, Sooyoung (Department of Computer Engineering, Handong Global University)

Publication Information

Journal of the Korea Institute of Information and Communication Engineering / v.18, no.3, 2014 , pp. 625-630 More about this Journal

Abstract

The amount of information is increasing rapidly with the development of the internet and the computer. Since these enormous information is managed by the document forms, it is necessary to search and process them efficiently. The document clustering technique which clusters the related documents through the similarity between the documents help to classify, search, and process the large amount of documents automatically. This paper proposes a method to find the initial seed points through principal component analysis when the documents represented by vectors in the feature vector space are clustered by K-means algorithm in order to increase clustering performance. The experiment shows that our method has a better performance than the traditional K-means algorithm.

Keywords

Document Clustering; K-means algorithm; PCA;

Citations & Related Records

Reference

1	C. Park, Y. Kim, J. Kim, J. Song, and H. Choi, Data Mining using R, Kyowoosa, 2011.
2	H. Park, and K. Lee, Pattern Recognition and Machine Learning from Basic to Application, Leehan Pub., 2011.
3	L. Oh, Pattern Recognition, Kyobo Book Centre, 2010.
4	S. Park, D. An, "Document Clustering Method using PCA and Fuzzy Association," Journal of Korea Information Processing Society B, 2010.
5	C. Lee, M. Kim, K. Lee, G. Lee, H. Park, "Document Thematic words Extraction using Principal Component Analysis," Journal of the Korea Society of Computer and Information B, 2002.
6	C. Lee, M. Kim, J. Paik, H. Park, "Text Summarization using PCA and SVD," Journal of Korea Information Processing Society B, 2003.
7	S. Park, J. Lee, "Topic-basied Multi-document Summarization Using Non-negative Matrix Factorization and K-means," Journal of the Korea Society of Computer and Information B, 2008.
8	S. Park, D. U. An, B. R. Char, and C. W. Kim, "Document Clustering with Cluster Refinement and Non-negative Matrix Factorization," In Proceeding of ICONIP'09, 2009.
9	S. Osinski and D. Weiss, "Conceptua Clustering using lingo algorithm: Evaluation on open directory project data," in Proc. IIPWM04, 2004.
10	The Porter Stemming Algorithm. Available: http://tartarus.org/-martin/PorterStemmer/
11	B. Lee, Information Retrieval, Green Pub. 2012.
12	http://qwone.com/-jason/20Newsgroups/

1	(2014) International journal of advanced smart convergence A Study on Efficient Memory Management Using Machine Learning Algorithm / 6 (1) , 39
1	(2014) 정보관리학회지 단어 임베딩(Word Embedding) 기법을 적용한 키워드 중심의 사회적 이슈 도출 연구: 장애인 관련 뉴스 기사를 중심으로 / 35 (1) , 231
8	(2014) 한국정보통신학회논문지 잠재 의미 분석을 적용한 유사 특허 검색 서비스 시스템 / 22 (8) , 1049
3	(2014) 한국항만경제학회지 K-Means 군집모형과 계층적 군집(교차효율성 메트릭스에 의한 평균연결법, Ward법)모형 및 혼합모형을 이용한 컨테이너항만의 클러스터링 측정에 대한 실증적 비교 및 검증에 관한 연구 / 34 (3) , 17

KSCI

Document Clustering Technique by K-means Algorithm and PCA 주성분 분석과 k 평균 알고리즘을 이용한 문서군집 방법

Document Clustering Technique by K-means Algorithm and PCA