• Title/Summary/Keyword: ??치

Search Result 27,144, Processing Time 0.045 seconds

Discovery of Interesting Knowlege using Concept Hierarchy (개념 계층 이용 흥미로운 부분 데이터의 탐색)

  • 홍정희;김성민;남도원;이동하;이전영
    • Proceedings of the Korea Inteligent Information System Society Conference
    • /
    • 2000.04a
    • /
    • pp.261-270
    • /
    • 2000
  • 개념 계층(Concept Hierarchy)은 데이터베이스 분야에서 사용되는 대표적인 배경 지식(Background Knowledge)으로써, 데이터베이스에 내재되어 있는 구조적인 정보, 데이터의 분포, 영역전문가(Domain Expert)에 의해 주어지는 외부 지식 등이 반영되어 있다. 개념 계층의 특성상 부모(parent)-자식(child) 관계가 있는 두 노드가 있을 때, 한 노드의 값으로부터 다른 노드의 값을 추정할 수 있다. 이 추정된 값을 기대치라고 하고, 한 노드의 값으로부터 추정된 기대치와 실제치가 상당히 상이한 값을 보이는 노드가 있을 때, 이를 흥미롭다(interesting)라고 할 수 있다. 그러나 아직까지 개념계층상에서의 흥미로운 부분 탐색에 대한 연구가 없었으며, 흥미로움(interestingness)의 척도(measurement)에 대한 연구로서는 신뢰도(confidence), 리프트(lift), 컨빅션(conviction)등이 있다. 그러나 이런 흥미도의 척도에 관한 연구도 연관규칙에 한정되어 이루어졌으므로 개념계층상의 데이터에 적용하기 위해서는 약간의 수정 및 새로운 정의가 필요하다. 본 논문에서는 데이터의 특성에 따른 개념계층이 존재할 때, 이를 이용하여 기대치와 실제치가 상이한 흥미로운 부분을 발견하고자 하며, 이를 위하여 개념계층이 존재할 때, 이를 이용하여 기대치와 실제치가 상이한 흥미로운 부분을 발견하고자 하며, 이를 위하여 개념계층상에서의 흥미도의 척도를 제안하고 흥미로운 부분을 탐색하는 방법을 기술하고자 한다. 또한 데이터마이닝의 결과인 연관규칙을 개념계층에 적용하여 연관규칙을 통해 얻어질 수 있는 기대치를, 지지도(support), 신뢰도(confidence), 리프트(lift), 컨빅션(conviction)등의 관계를 통해 다양한 방법으로 모색해본다. 이 연구에서 제안하는 이러한 개념계층상의 흥미로운 부분의 탐색은, 전자 상거래에서의 CRM(Customer Relationship Management)나 틈새시장(niche market) 마케팅 등에 적용가능하리라 여겨진다.

  • PDF

Robust tests for heteroscedasticity using outlier detection methods (이상치 탐지법을 이용한 강건 이분산 검정)

  • Seo, Han Son;Yoon, Min
    • The Korean Journal of Applied Statistics
    • /
    • v.29 no.3
    • /
    • pp.399-408
    • /
    • 2016
  • There is a need to detect heteroscedasticity in a regression analysis; however, it invalidates the standard inference procedure. The diagnostics on heteroscedasticity may be distorted when both outliers and heteroscedasticity exist. Available heteroscedasticity detection methods in the presence of outliers usually use robust estimators or separating outliers from the data. Several approaches have been suggested to identify outliers in the heteroscedasticity problem. In this article conventional tests on heteroscedasticity are modified by using a sequential outlier detection methods to separate outliers from contaminated data. The performance of the proposed method is compared with original tests by a Monte Carlo study and examples.

Outlier Detection Using Dynamic Plots (동적 그림을 이용한 이상치 검색)

  • Ahn, Byung-Jin;Seo, Han-Son
    • The Korean Journal of Applied Statistics
    • /
    • v.24 no.5
    • /
    • pp.979-986
    • /
    • 2011
  • A linear regression method is commonly used to analyze data because of its simplicity and applicability; however, it is well known that data may contain some outliers and influential cases that may have a harmful effect on a statistical analysis. Thus detection and examination of outliers or influential cases are important parts of data analysis. In detecting multiple outliers, masking effects usually occur and make it difficult to identify the true outliers. We propose to use dynamic plots as a method resistant to masking effect. The procedure using dynamic plots is useful to find appropriate basic sets with which a dependent outliers detection method start and detect a true outliers set. Examples are given to demonstrate the effectiveness of the suggested idea.

A Design of a Ternary Storage Elements Using CMOS Ternary Logic Gates (CMOS 3치 논리 게이트를 이용한 3치 저장 소자 설계)

  • Yoon, Byoung-Hee;Byun, Gi-Young;Kim, Heung-Soo
    • Journal of IKEEE
    • /
    • v.8 no.1 s.14
    • /
    • pp.47-53
    • /
    • 2004
  • We present the design of ternary flip-flop which is based on ternary logic so as to process ternary data. These flip-flops are composed with ternary voltage mode NMAX, NMIN, INVERTER gates. These logic gate circuits are designed using CMOS and obtained the characteristics of a lower voltage, lower power consumption as compared to other gates. These circuits have been simulated with the electrical parameters of a standard 0.35um CMOS technology and 3.3Volts supply voltage. The architecture of proposed ternary flip-flop is highly modular and well suited for VLSI implementation, only using ternary gates.

  • PDF

A Performance Comparison of Machine Learning Library based on Apache Spark for Real-time Data Processing (실시간 데이터 처리를 위한 아파치 스파크 기반 기계 학습 라이브러리 성능 비교)

  • Song, Jun-Seok;Kim, Sang-Young;Song, Byung-Hoo;Kim, Kyung-Tae;Youn, Hee-Yong
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2017.01a
    • /
    • pp.15-16
    • /
    • 2017
  • IoT 시대가 도래함에 따라 실시간으로 대규모 데이터가 발생하고 있으며 이를 효율적으로 처리하고 활용하기 위한 분산 처리 및 기계 학습에 대한 관심이 높아지고 있다. 아파치 스파크는 RDD 기반의 인 메모리 처리 방식을 지원하는 분산 처리 플랫폼으로 다양한 기계 학습 라이브러리와의 연동을 지원하여 최근 차세대 빅 데이터 분석 엔진으로 주목받고 있다. 본 논문에서는 아파치 스파크 기반 기계 학습 라이브러리 성능 비교를 통해 아파치 스파크와 연동 가능한 기계 학습라이브러리인 MLlib와 아파치 머하웃, SparkR의 데이터 처리 성능을 비교한다. 이를 위해, 대표적인 기계 학습 알고리즘인 나이브 베이즈 알고리즘을 사용했으며 학습 시간 및 예측 시간을 비교하여 아파치 스파크 기반에서 실시간 데이터 처리에 적합한 기계 학습 라이브러리를 확인한다.

  • PDF

Outlier Detection Using Support Vector Machines (서포트벡터 기계를 이용한 이상치 진단)

  • Seo, Han-Son;Yoon, Min
    • Communications for Statistical Applications and Methods
    • /
    • v.18 no.2
    • /
    • pp.171-177
    • /
    • 2011
  • In order to construct approximation functions for real data, it is necessary to remove the outliers from the measured raw data before constructing the model. Conventionally, visualization and maximum residual error have been used for outlier detection, but they often fail to detect outliers for nonlinear functions with multidimensional input. Although the standard support vector regression based outlier detection methods for nonlinear function with multidimensional input have achieved good performance, they have practical issues in computational cost and parameter adjustments. In this paper we propose a practical approach to outlier detection using support vector regression that reduces computational time and defines outlier threshold suitably. We apply this approach to real data examples for validity.

Probability Characteristics of Probable Rainfall and Recorded Maximum Rainfall in Korea. (한국주요지점에 대한 확률강우량과 관측최대강우량의 확률분석)

  • Jeong, Mahn;Lee, Jong-Kyu
    • Water for future
    • /
    • v.14 no.3
    • /
    • pp.47-54
    • /
    • 1981
  • The characteristics of point rainfall for three different durations in Seoul Pusan Taegu and Gwangju have been analysed by the probabilistic ainfall method and the M-year maximum rainfall method. The probabilities that the T-year probabilistic rainfall did not occur during the observation period, compared with the values obtained from the observed data. were smaller than the theoretical values. The averages of the probabilities that the M-year maximum-ten-minute rainfall did not occur in the consequent N-years were larger than the theoretical values, the M-year maximumone hour rainfall were smaller than the theoretical ones, and the M-year maximum daily rainfall nearly agreed with them, and while those of Japan were smaller than the theoretical values. It is recommended from the results that the recorded maximum value should be used as a design value rather than the probabilistic rainfall.

  • PDF

Study on Material Properties of Composite Materials using Finite Element Method (유한요소법을 이용한 복합재의 물성치 도출에 대한 연구)

  • Jung, Chul-Gyun;Kim, Sung-Uk
    • Journal of the Computational Structural Engineering Institute of Korea
    • /
    • v.29 no.1
    • /
    • pp.61-65
    • /
    • 2016
  • Composites are materials that are widely used in industries such as automobile and aircraft. The composite material is required as a material for using in a high temperature environment as well as acting as a high pressure environment like the nozzle part of the ship. It is important to know the properties of composites. Result values obtained substituting the properties of matrix and fiber numerically have an large error compared with experimental value. In this study we utilize CASADsolver EDISON program for using Finite Element Method. Properties by substituting the fiber and Matrix properties of the composite material properties were compared with those measured in the experiment and calculated by the empirical properties.

Flame detection algorithm using adaptive threshold in thermal video (적응 문턱치를 이용한 열영상 화염 검출 알고리즘)

  • Jeong, Soo-Young;Kim, Won-Ho
    • Journal of Satellite, Information and Communications
    • /
    • v.9 no.4
    • /
    • pp.91-96
    • /
    • 2014
  • This paper proposed an adaptive threshold method for detecting flame candidate regions in a infrared image and it adapts according to the contrast and intensity changes in the image. Conventional flame detection systems uses fixed threshold method since surveillance environment does not change, once the system installed. But it needs a adaptive threshold method as requirements of surveillance system has changed. The proposed adaptive threshold algorithm uses the dynamic behavior of flame as featured parameter. The test result is analysed by comparing test result of proposed adaptive threshold algorithm and conventional fixed threshold method. The analysed data shows, the proposed method has 91.42% of correct detection rate and false detection is reduced by 20% comparing to the conventional method.

Outlier detection and time series modelling in the stationary time series (정상 시계열에서의 이상치 발견과 시계열 모형구축)

  • 이종협;최기헌
    • The Korean Journal of Applied Statistics
    • /
    • v.5 no.2
    • /
    • pp.139-156
    • /
    • 1992
  • Recently several authors have introduced iterative methods for detecting time series outliers. Most of these methods are developed under the assumption that an underlying outlier-free model is known or can be identified. Since outliers can distort model identification or even make it impossible, we propose procedure begins with a descriptive data analysis of a time series using distance measures between two observations. Properties of the proposed test statistic are presented. To distinguish the type of an outlier are used transfer function models. An empirical example is given to illustrate the time series modeling procedure.

  • PDF