• Title/Summary/Keyword: Term Clustering

Search Result 177, Processing Time 0.024 seconds

An Experimental Study on Selecting Association Terms Using Text Mining Techniques (텍스트 마이닝 기법을 이용한 연관용어 선정에 관한 실험적 연구)

  • Kim, Su-Yeon;Chung, Young-Mee
    • Journal of the Korean Society for information Management
    • /
    • v.23 no.3 s.61
    • /
    • pp.147-165
    • /
    • 2006
  • In this study, experiments for selection of association terms were conducted in order to discover the optimum method in selecting additional terms that are related to an initial query term. Association term sets were generated by using support, confidence, and lift measures of the Apriori algorithm, and also by using the similarity measures such as GSS, Jaccard coefficient, cosine coefficient, and Sokal & Sneath 5, and mutual information. In performance evaluation of term selection methods, precision of association terms as well as the overlap ratio of association terms and relevant documents' indexing terms were used. It was found that Apriori algorithm and GSS achieved the highest level of performances.

A Text Summarization Model Based on Sentence Clustering (문장 클러스터링에 기반한 자동요약 모형)

  • 정영미;최상희
    • Journal of the Korean Society for information Management
    • /
    • v.18 no.3
    • /
    • pp.159-178
    • /
    • 2001
  • This paper presents an automatic text summarization model which selects representative sentences from sentence clusters to create a summary. Summary generation experiments were performed on two sets of test documents after learning the optimum environment from a training set. Centroid clustering method turned out to be the most effective in clustering sentences, and sentence weight was found more effective than the similarity value between sentence and cluster centroid vectors in selecting a representative sentence from each cluster. The result of experiments also proves that inverse sentence weight as well as title word weight for terms and location weight for sentences are effective in improving the performance of summarization.

  • PDF

A Design of an Improved Linguistic Model based on Information Granules (정보 입자에 근거한 개선된 언어적인 모델의 설계)

  • Han, Yun-Hee;Kwak, Keun-Chang
    • Journal of the Institute of Electronics Engineers of Korea CI
    • /
    • v.47 no.3
    • /
    • pp.76-82
    • /
    • 2010
  • In this paper, we develop Linguistic Model (LM) based on information granules as a systematic approach to generating fuzzy if-then rules from a given input-output data. The LM introduced by Pedrycz is performed by fuzzy information granulation obtained from Context-based Fuzzy Clustering(CFC). This clustering estimates clusters by preserving the homogeneity of the clustered patterns associated with the input and output data. Although the effectiveness of LM has been demonstrated in the previous works, it needs to improve in the sense of performance. Therefore, we focus on the automatic generation of linguistic contexts, addition of bias term, and the transformed form of consequent parameter to improve both approximation and generalization capability of the conventional LM. The experimental results revealed that the improved LM yielded a better performance in comparison with LM and the conventional works for automobile MPG(miles per gallon) predication and Boston housing data.

k-Bitmap Clustering Method for XML Data based on Relational DBMS (관계형 DBMS 기반의 XML 데이터를 위한 k-비트맵 클러스터링 기법)

  • Lee, Bum-Suk;Hwang, Byung-Yeon
    • The KIPS Transactions:PartD
    • /
    • v.16D no.6
    • /
    • pp.845-850
    • /
    • 2009
  • Use of XML data has been increased with growth of Web 2.0 environment. XML is recognized its advantages by using based technology of RSS or ATOM for transferring information from blogs and news feed. Bitmap clustering is a method to keep index in main memory based on Relational DBMS, and which performed better than the other XML indexing methods during the evaluation. Existing method generates too many clusters, and it causes deterioration of result of searching quality. This paper proposes k-Bitmap clustering method that can generate user defined k clusters to solve above-mentioned problem. The proposed method also keeps additional inverted index for searching excluded terms from representative bits of k-Bitmap. We performed evaluation and the result shows that the users can control the number of clusters. Also our method has high recall value in single term search, and it guarantees the searching result includes all related documents for its query with keeping two indices.

Analysis of Departing Passengers' Dwell Time using Clustering Techniques (클러스터링 기법을 활용한 출발 여객 체류 시간 분석)

  • An, Deok-bae;Kim, Hui-yang;Baik, Ho-jong
    • Journal of Advanced Navigation Technology
    • /
    • v.23 no.5
    • /
    • pp.380-385
    • /
    • 2019
  • This paper is concerned with departure passengers' dwell time analysis using real system data. Previous researches emphasize the importance of dwell time analysis from perspective of airport terminal planning and non-aeronautical revenue. However, short-term airport operation using passengers' dwell time is considered impossible due to absence of passengers' behavior data. Recently, in accordance with the wave of smart airport, world leading airports are systematically collecting passenger data. So there is high possibility of analyzing passengers' dwell time with the data stacked in the airport database. We conducted dwell time analysis using data from Incheon Int'l airport. In order to handle passenger data, we adapted clustering algorithm which is one of data mining techniques. As a clustering result, passengers are divided into 3 clusters. One is the cluster for passengers whose dwell time is relatively short and who tend to spend longer time in the airside. Another is the cluster for passengers who have near 3 hours dwell time. The other is the cluster for passengers whose total dwell time is extremely long.

Analytical Study of Fuzzy Clustering Technique for Automatic Term Classification (용어 자동분류를 위한 퍼지 클러스터링 기법 분석)

  • 한승희
    • Proceedings of the Korean Society for Information Management Conference
    • /
    • 2003.08a
    • /
    • pp.95-103
    • /
    • 2003
  • 목차 및 권말색인과 같은 인쇄형태의 정보내용에 대한 구조화된 접근방식에서 착안하여 전자 문서의 내용에 대한 새로운 형태의 접근방식을 개발할 수 있는데, 이를 위한 방안으로 용어 자동분류 기법이 있다. 본 연구에서는 용어의 의미모호성 문제를 해결하는 동시에 용어간 계층관계 표현이 가능한 자동분류 기법으로 퍼지 클러스터링 기법을 제안하고, 대표적인 퍼지 클러스터링 알고리즘인 퍼지 c-means 기법에 대해 분석하고자 한다.

  • PDF

Web Document Clustering Using Statistical Techniques & Tag Information on the Specific-Domain Web site (전문 웹 사이트에서의 통계적 기법과 태그 정보를 이용한 문서 분류)

  • 조은휘;변영태
    • Proceedings of the Korea Inteligent Information System Society Conference
    • /
    • 2002.11a
    • /
    • pp.297-302
    • /
    • 2002
  • 특정 영역에 대해 사용자에게 관련 정보를 제공하는 서비스를 위해 정보 에이전트를 개발하고 있다. 이 시스템은 웹 상에서 문서를 수집해 오는데 특정 영역과 관련한 지식베이스를 토대로 하고 있는데, 이들 중 몇몇 전문 사이트 내의 정보가 많이 포함되어 있음을 볼 수 있다. 그러므로 전문 사이트 내의 관련 문서 수집은 중요한 의의가 있다. 본 논문에서는 이들 전문 사이트 내의 전문 문서 수집을 위해 문서간의 유사성을 토대로 클러스터링 한다. 즉, 문서내의 텀(term)과 HTML 태그(tag), 지식베이스의 WordNet 계층구조를 data로 하고 SVD(Singular Value Decomposition)을 사용하여 문서간의 관계를 밝혀내었다.

  • PDF

Speed Control of BLDC Motor Drive Using an Adaptive Fuzzy P+ID Controller (적응 퍼지 P+ID 제어기를 이용한 BLDC 전동기의 속도제어)

  • Kwon, Chung-Jin;Han, Woo-Yang;Sin, Dong-Yang;Kim, Sung-Joong
    • Proceedings of the KIEE Conference
    • /
    • 2002.07b
    • /
    • pp.1172-1174
    • /
    • 2002
  • An adaptive fuzzy P + ID controller for variable speed operation of BLDC motor drives is presented in this paper. Generally, a conventional PID controller is most widely used in industry due to its simple control structure and ease of design. However, the PID controller suffers from the electrical machine parameter variations and disturbances. To improve the tracking performance for parameter and load variations, the controller proposed in this paper is constructed by using an adaptive fuzzy logic controller in place of the proportional term in a conventional PID controller. For implementing this controller, only one additional parameter has to be adjusted in comparison with the PID controller. An adaptive fuzzy controller applied to proportional term to achieve robustness against parameter variations has simple structure and computational simplicity. The controller based on optimal fuzzy logic controller has an self-tuning characteristics with clustering. Computer simulation results show the usefulness of the proposed controller.

  • PDF

electrical Damage of Metallized Film Capacitors (필름 Capacitor의 전기적Damage에 관한 연구)

  • ;Chathan M. Cooke
    • The Transactions of the Korean Institute of Electrical Engineers
    • /
    • v.40 no.6
    • /
    • pp.574-581
    • /
    • 1991
  • Damage in film capacitors has been investigated, using FTIR and ESCA, aiming to elucidate the nature of electrode removal and the possibility of base films to be damaged. Also, tests were conducted to investigate the effect of a long-term thermal aging at elevated temperatures. Unsuccessful clearing or grape-clustering processes can induce a long-term degradation which involves the chemical and morphological changes. Major changes are the oxidation and the decrease in surface crystallinity possibly arising from the corona discharge. An immediate deterioration of BOPP film may occur when the air entrapped between the film layers induces an extensive autocatalytic oxidative degradation. This type of immediate damage may result in a premature failure at an early stage of qualification test. As far as the nature of electrode removal is concerned, a permanent removal of electrode materials was observed in the main erosion area.

  • PDF

Development of A Web Mining System Based On Document Similarity (문서 유사도 기반의 웹 마이닝 시스템 개발)

  • 이강찬;민재홍;박기식;임동순;우훈식
    • The Journal of Society for e-Business Studies
    • /
    • v.7 no.1
    • /
    • pp.75-86
    • /
    • 2002
  • In this study, we proposed design issues and structure of a web mining system and develop a system for the purpose of knowledge integration under world wide web environments resulted from our developing experiences. The developed system consists of three main functions: 1) gathering documents utilizing a search agent; 2) determining similarity coefficients between any two documents from term frequencies; 3) clustering documents based on similarity coefficients. It is believed that the developed system can be utilized for discovery of knowledge in relatively narrow domains such as news classification, index term generation in knowledge management.

  • PDF