한국멀티미디어학회논문지 (Journal of Korea Multimedia Society)
- 제11권12호
- /
- Pages.1749-1757
- /
- 2008
- /
- 1229-7771(pISSN)
- /
- 2384-0102(eISSN)
Single Pass Algorithm for Text Clustering by Encoding Documents into Tables
초록
This research proposes a modified version of single pass algorithm specialized for text clustering. Encoding documents into numerical vectors for using the traditional version of single pass algorithm causes the two main problems: huge dimensionality and sparse distribution. Therefore, in order to address the two problems, this research modifies the single pass algorithm into its version where documents are encoded into not numerical vectors but other forms. In the proposed version, documents are mapped into tables and the operation on two tables is defined for using the single pass algorithm. The goal of this research is to improve the performance of single pass algorithm for text clustering by modifying it into the specialized version.