Search | Korea Science

Feature Selection by Genetic Algorithm and Information Theory (유전자 알고리즘과 정보이론을 이용한 속성선택)

Jo, Jae-Hun
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 2007.11a
- /
- pp.108-111
- /
- 2007
속성선택(Feature Selection)은 패턴분류 문제에서 분류기들의 성능을 향상시킬 수 있는 중요한 부분으로 다양한 기법들이 연구되어지고 있다. 특히, 많은 변수와 속성들을 가지는 데이터를 패턴분류 하는 과정에서 주요 속성부분집합을 추출하여 이용함으로써 분류기의 연산속도 및 정확도를 향상시킬 수 있다. 본 논문에서는 유전자 알고리즘과 정보이론의 상호정보량을 이용하여 속성선택을 하는 기법을 제안하였다. 제안된 기법의 성능을 평가하기 위하여 패턴분류 문제에 적용하고 그 성능이 우수함을 확인하였다.
PDF

Application of Cluster Analysis using Mutual Information (상호정보량 기법을 이용한 군집분석의 적용성 연구)

Jung, Young-Hun;Kim, Wan-Su;Jeong, Chang-Sam;Heo, Jun-Haeng
- Proceedings of the Korea Water Resources Association Conference
- /
- 2011.05a
- /
- pp.414-414
- /
- 2011
우리나라 뿐만 아니라 전 세계적으로 기후변화로 인한 집중호우, 폭설 등이 빈번하게 일어나고 있으며 수공구조물 설계에 필요한 확률강우량도 증가하고 있다. 확률강우량을 산정하는 빈도해석의 경우 지점빈도해석의 문제점을 보완한 지역빈도해석에 대한 연구가 꾸준히 진행되고 있다. 지역빈도해석을 적용하기 위해서는 수문학적 동질성을 가지는 지역 구분이 무엇보다 중요하다. 군집 분석은 개체들이 지니고 있는 다양한 속성의 유사성을 동질적인 집단으로 군집화하는 방법을 말한다. 군집분석의 기본원리는 분석하고자 하는 여러 특성등을 유사성(similaruty) 거리(distance)로 환산하고 거리가 상대적으로 가까운 개체들을 동질적으로 군집화하는 것이다. 군집분석을 적용하기 위해서는 기상학적 인자와 지형학적 인자를 이용하여 군집분석을 실시한다. 군집분석을 실시할 때 가장 중요한 것은 입력변수의 선택으로 입력 변수의 적절한 선택이 결과값에 큰 영향을 준다. 상호정보량(Mutual Information, MI) 기법은 두 무작위 변수간의 관련성을 측정하는 방법이며 (Cover and Tomas, 2006), 두 변수간의 독립성 구조에 관한 가정이 없고 데이터 변형이나 잡음(noise)에 대한 영향이 적어 다른 기법보다 신뢰도가 높다고 알려져 있다(Peng et al., 2005). 본 연구에서는 상호정보량 기법을 이용하여 군집된 지점들의 종속성과 독립성의 관계를 정량적으로 산정하여 비교하고자 한다.
PDF

Application on Prediction of Stream Flow using Artificial Neural Network with Mutual Information and Wavelet Transform (상호정보량기법과 웨이블렛변환을 적용한 인공신경망의 하천유량 예측 활용)

Ryu, Yong-Jun;Jung, Yong-Hun;Shin, Ju-Young;Heo, Jun-Haeng
- Proceedings of the Korea Water Resources Association Conference
- /
- 2012.05a
- /
- pp.116-116
- /
- 2012
하천유역 내의 인자를 이용하여 댐의 하천유량(stream flow)을 예측하는 일은 수문특성의 연구와 자연재해에 대한 대비 및 수공구조물과 방재시설의 설계 시 중요한 역할을 한다. 이러한 연구는 과거부터 활발히 이루어졌으며, 아직도 보다 높은 정확도의 결과를 얻기 위해 많은 연구들이 이루어지고 있다. 특히 기존의 유역 내 자료를 통해 비선형적 모델인 인공신경망(artificial neural network)을 이용한 하천유량을 예측하는 연구 역시 활발히 이루어지고 있다. 본 연구의 목적은 여러 유역인자들 중 하천유량에 가장 영향을 미치는 변수를 추출하고 보다 정확한 예측모델을 구축하는 것이다. 기존의 입력자료 선정기법중의 하나인 상호정보량(mutual information)과 수문기상자료의 비선형 동역학적 성분을 추출하는 웨이블렛 변환(wavelet transform)을 사용하여 인공신경망에 적용시켰다. 인공신경망을 적용하는 경우, 수문자료에 있어서 변수의 선택과 자료의 상태가 강우예측의 결과에 큰 영향을 미친다. 이러한 변수의 선택에 있어서 상호정보량을 바탕으로 한 인공신경망 입력변수 선택기법이 많이 사용되고 있다. 일반적으로 시계열자료는 경향성(trend), 주기성(periodicity) 및 추계학적 성분(stochastic component)의 선형조합으로 가정될 수 있으며, 특히 경향성과 주기성은 시계열 모형을 위해 제거되어야 할 결정론적 성분으로 취급한다. 즉. 수문 기상자료에 포함되어 있는 경향성과 주기성과 같은 비선형 동역학적 잡음(nonlinear dynamical noise)을 제거하고 입력자료의 카오스적 거동을 보이는 성분을 분리하기 위해 웨이블렛 변환을 사용하였다. 대상유역은 한강 유역에 포함되어 있는 충주댐으로 선택하였다. 유역 내 다양한 인자들과 하천유량사이의 상호정보량을 구해 영향력이 가장 큰 변수를 추출하고, 그 자료를 웨이블렛 변환을 적용하여 인공신경망의 입력자료로 사용하였다. 본 논문에서는 위와 같은 과정을 이용해 추정한 하천유량 결과와 기존의 방법인 상호정보량을 이용해 인공신경망을 적용한 결과를 실제자료와 비교하였다.
PDF

Calibration of Real Time Rainfall Data Using Mutual Information and Artificial Neural Network (상호정보량 기법과 인공신경망을 이용한 실시간 강우 자료 보정)

Sung, Kyung-Min;Goo, Yeo-Joo;Kim, Tae-Soon;Heo, Jun-Haeng
- Proceedings of the Korea Water Resources Association Conference
- /
- 2010.05a
- /
- pp.1269-1273
- /
- 2010
이러한 강우자료의 결측값이나 오자료를 보정하는 것은 그 유역의 정확한 수문학적 특성 파악 및 안전한 수공구조물의 설계에 영향을 미치게 되므로 매우 중요하다고 할 수 있다. 최근 이러한 강우자료를 비선형적 모델인 인공신경망(Artificial Neural Network)을 이용하여 보정하는 연구가 활발히 진행되고 있다(오재우 등, 2008). 그러나 이러한 인공신경망을 적용하는 경우, 선택한 신경망 구조의 형태와 학습(training)을 위해 사용되는 자료가 전체 자료의 특성을 반영하고 있는 정도에 따라 정확도에 차이를 보인다(한광희 등, 2010). 따라서 자료보정을 위한 입력 자료의 선택은 인공신경망을 이용한 결측치 보정의 중요한 과정이다. 본 연구에서는 이러한 입력 자료의 선택을 위한 여러 가지 기법 중 입력 변수간의 상호정보량 (Mutual Information)을 이용한 방법을 적용하여 대상 결측 지점을 보정할 강우지점을 선별한 후 선택된 지점만으로 인공신경망을 구성하여 강우자료를 보정하고 주변 자료를 모두 이용한 결과와 상관성분석으로 얻어진 결과와 비교하였다.
PDF

Feature Selection Method by Information Theory and Particle S warm Optimization (상호정보량과 Binary Particle Swarm Optimization을 이용한 속성선택 기법)

Cho, Jae-Hoon;Lee, Dae-Jong;Song, Chang-Kyu;Chun, Myung-Geun
- Journal of the Korean Institute of Intelligent Systems
- /
- v.19 no.2
- /
- pp.191-196
- /
- 2009
In this paper, we proposed a feature selection method using Binary Particle Swarm Optimization(BPSO) and Mutual information. This proposed method consists of the feature selection part for selecting candidate feature subset by mutual information and the optimal feature selection part for choosing optimal feature subset by BPSO in the candidate feature subsets. In the candidate feature selection part, we computed the mutual information of all features, respectively and selected a candidate feature subset by the ranking of mutual information. In the optimal feature selection part, optimal feature subset can be found by BPSO in the candidate feature subset. In the BPSO process, we used multi-object function to optimize both accuracy of classifier and selected feature subset size. DNA expression dataset are used for estimating the performance of the proposed method. Experimental results show that this method can achieve better performance for pattern recognition problems than conventional ones.
https://doi.org/10.5391/JKIIS.2009.19.2.191 인용 PDF KSCI

Classification of Hyperspectral Images Using Spectral Mutual Information (분광 상호정보를 이용한 하이퍼스펙트럴 영상분류)

Byun, Young-Gi;Eo, Yang-Dam;Yu, Ki-Yun
- Journal of Korean Society for Geospatial Information Science
- /
- v.15 no.3
- /
- pp.33-39
- /
- 2007
Hyperspectral remote sensing data contain plenty of information about objects, which makes object classification more precise. In this paper, we proposed a new spectral similarity measure, called Spectral Mutual Information (SMI) for hyperspectral image classification problem. It is derived from the concept of mutual information arising in information theory and can be used to measure the statistical dependency between spectra. SMI views each pixel spectrum as a random variable and classifies image by measuring the similarity between two spectra form analogy mutual information. The proposed SMI was tested to evaluate its effectiveness. The evaluation was done by comparing the results of preexisting classification method (SAM, SSV). The evaluation results showed the proposed approach has a good potential in the classification of hyperspectral images.
PDF

An Adaptive Sequential Prefetching using Traffic Information in Shared-Memory Multiprocessors (공유메모리 다중처리기에서 상호연결망의 통신량을 고려하는 선인출 기법)

박정우;손영철;정한조;맹승렬
- Proceedings of the Korean Information Science Society Conference
- /
- 2000.04a
- /
- pp.633-635
- /
- 2000
상호연결망을 기반으로 하는 공유메모리 다중처리기의 성능은 공유메모리 접근 속도에 많은 영향을 받는다. 선인출 기법은 프로세서의 계산과 데이터의 접근을 중첩시켜 메모리의 접근 속도를 줄인다. 기존의 선인출 기법들은 캐쉬미스 양을 줄이는 것만을 생각하여 상호연결망의 상황을 고려하지 않은 문제점이 있다. 본 논문에서는 응답이 늦은 선인출 이용하여 선인출 양을 조절함으로써 상호연결망의 경쟁을 줄이는 새로운 선인출 기법을 제안하고 프로그램 구동 모의실험을 통해 기존의 선인출 기법[1]에 비해 더 좋은 성능을 나타냄을 보인다.
PDF

Double Clustering of Gene Expression Data Based on the Information Bottleneck Method (정보병목기법에 기반한 유전자 발현 데이터의 이중 클러스터링)

김병희;황규백;장정호;장병탁
- Proceedings of the Korean Information Science Society Conference
- /
- 2003.04c
- /
- pp.362-364
- /
- 2003
기능 유전체학에서 클러스터링 기법은 고차원의 마이크로 어레이 데이터 분석을 위한 주된 도구 중의 하나이다. 본 논문에서는 정보병목(information bottleneck)기법 기반의 이중 클러스터링에 의한, 유전자 발현 데이터의 계층적 병합방식 클러스터링 기법을 제안한다. 정보병목기법은, 두 랜덤변수의 결합확률분포가 주어진 경우 두 변수의 상호 정보량을 최대한 보존하면서 한 변수를 압축하는 기법이며, 두 변수를 차례로 압축하는 것이 이중 클러스터링이다. 실제 마이크로 어레이 데이터인 NC160 데이터(암세포 내 유전자 발현 데이터)에 대한 실험에서, 먼저 유전자를 그 발현패턴에 따라 클러스터링 한 후 이를 이용하여 표본들을 클러스터링하고 그 성능을 다각도로 분석하였다. 상호 정보량과 유전자 및 표본 클러스터 수와 엔트로피 척도에 의한 성능을 검토해 본 결과, 표본이 추출 조직에 따라 구분 가능할 것이라는 가정을 검증할 수 있었으며, 적절한 클러스터의 수를 결정할 수 있는 임계점의 기준을 설정할 수 있었다.
PDF

Input Variables Selection of Artificial Neural Network Using Mutual Information (상호정보량 기법을 적용한 인공신경망 입력자료의 선정)

Han, Kwang-Hee;Ryu, Yong-Jun;Kim, Tae-Soon;Heo, Jun-Haeng
- Journal of Korea Water Resources Association
- /
- v.43 no.1
- /
- pp.81-94
- /
- 2010
Input variable selection is one of the various techniques for improving the performance of artificial neural network. In this study, mutual information is applied for input variable selection technique instead of correlation coefficient that is widely used. Among 152 variables of RDAPS (Regional Data Assimilation and Prediction System) output results, input variables for artificial neural network are chosen by computing mutual information between rainfall records and RDAPS' variables. At first the rainfall forecast variable of RDAPS result, namely APCP, is included as input variable and the other input variables are selected according to the rank of mutual information and correlation coefficient. The input variables using mutual information are usually those variables about wind velocity such as D300, U925, etc. Several statistical error estimates show that the result from mutual information is generally more accurate than those from the previous research and correlation coefficient. In addition, the artificial neural network using input variables computed by mutual information can effectively reduce the relative errors corresponding to the high rainfall events.
https://doi.org/10.3741/JKWRA.2010.43.1.81 인용 PDF KSCI

Automatic Construction of Reduced Dimensional Cluster-based Keyword Association Networks using LSI (LSI를 이용한 차원 축소 클러스터 기반 키워드 연관망 자동 구축 기법)

Yoo, Han-mook;Kim, Han-joon;Chang, Jae-young
- Journal of KIISE
- /
- v.44 no.11
- /
- pp.1236-1243
- /
- 2017
In this paper, we propose a novel way of producing keyword networks, named LSI-based ClusterTextRank, which extracts significant key words from a set of clusters with a mutual information metric, and constructs an association network using latent semantic indexing (LSI). The proposed method reduces the dimension of documents through LSI, decomposes documents into multiple clusters through k-means clustering, and expresses the words within each cluster as a maximal spanning tree graph. The significant key words are identified by evaluating their mutual information within clusters. Then, the method calculates the similarities between the extracted key words using the term-concept matrix, and the results are represented as a keyword association network. To evaluate the performance of the proposed method, we used travel-related blog data and showed that the proposed method outperforms the existing TextRank algorithm by about 14% in terms of accuracy.
https://doi.org/10.5626/JOK.2017.44.11.1236 인용 KSCI

Search Result 159, Processing Time 0.03 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)