• 제목/요약/키워드: outlier identification

검색결과 36건 처리시간 0.021초

Identification of Incorrect Data Labels Using Conditional Outlier Detection

  • Hong, Charmgil
    • 한국멀티미디어학회논문지
    • /
    • 제23권8호
    • /
    • pp.915-926
    • /
    • 2020
  • Outlier detection methods help one to identify unusual instances in data that may correspond to erroneous, exceptional, or surprising events or behaviors. This work studies conditional outlier detection, a special instance of the outlier detection problem, in the context of incorrect data label identification. Unlike conventional (unconditional) outlier detection methods that seek abnormalities across all data attributes, conditional outlier detection assumes data are given in pairs of input (condition) and output (response or label). Accordingly, the goal of conditional outlier detection is to identify incorrect or unusual output assignments considering their input as condition. As a solution to conditional outlier detection, this paper proposes the ratio-based outlier scoring (ROS) approach and its variant. The propose solutions work by adopting conventional outlier scores and are able to apply them to identify conditional outliers in data. Experiments on synthetic and real-world image datasets are conducted to demonstrate the benefits and advantages of the proposed approaches.

Temporal and spatial outlier detection in wireless sensor networks

  • Nguyen, Hoc Thai;Thai, Nguyen Huu
    • ETRI Journal
    • /
    • 제41권4호
    • /
    • pp.437-451
    • /
    • 2019
  • Outlier detection techniques play an important role in enhancing the reliability of data communication in wireless sensor networks (WSNs). Considering the importance of outlier detection in WSNs, many outlier detection techniques have been proposed. Unfortunately, most of these techniques still have some potential limitations, that is, (a) high rate of false positives, (b) high time complexity, and (c) failure to detect outliers online. Moreover, these approaches mainly focus on either temporal outliers or spatial outliers. Therefore, this paper aims to introduce novel algorithms that successfully detect both temporal outliers and spatial outliers. Our contributions are twofold: (i) modifying the Hampel Identifier (HI) algorithm to achieve high accuracy identification rate in temporal outlier detection, (ii) combining the Gaussian process (GP) model and graph-based outlier detection technique to improve the performance of the algorithm in spatial outlier detection. The results demonstrate that our techniques outperform the state-of-the-art methods in terms of accuracy and work well with various data types.

A study On An Identification of Interactions In A Nonreplicated Two-Way Layout With $L_1$-Estimation

  • Lee, Ki-Hoon
    • Communications for Statistical Applications and Methods
    • /
    • 제7권1호
    • /
    • pp.119-128
    • /
    • 2000
  • This paper proposes a method for detecting interactions in a two-way layout with one observation per cell. The identification of interactions in the model is not clear for they are confounding with error terms. The $L_1$-Estimation is robust with respect to a y-direction outlier in linear model so we are able to estimate main effects without affection of interactions, If an observation is classified as an outlier we conclude it contains an interaction. An empirical study compared with a classical method is performed.

  • PDF

Outlier Tests in Sample Surveys

  • Namkyung, Pyong;Lee, Joon Suk
    • Communications for Statistical Applications and Methods
    • /
    • 제7권2호
    • /
    • pp.447-456
    • /
    • 2000
  • In this paper, we considered three methods for outlier identification sample surveys. First, we studied method of handling and adjusting outliers in normal population. Second, we studied existing methods using mean, maximum and minimum and proposed a test using of median which well reflects characteristic of data regardless of sampling distribution. Finally, we showed our test using median works better than Dixon and mean test through simulation.

  • PDF

Variable Selection and Outlier Detection for Automated K-means Clustering

  • Kim, Sung-Soo
    • Communications for Statistical Applications and Methods
    • /
    • 제22권1호
    • /
    • pp.55-67
    • /
    • 2015
  • An important problem in cluster analysis is the selection of variables that define cluster structure that also eliminate noisy variables that mask cluster structure; in addition, outlier detection is a fundamental task for cluster analysis. Here we provide an automated K-means clustering process combined with variable selection and outlier identification. The Automated K-means clustering procedure consists of three processes: (i) automatically calculating the cluster number and initial cluster center whenever a new variable is added, (ii) identifying outliers for each cluster depending on used variables, (iii) selecting variables defining cluster structure in a forward manner. To select variables, we applied VS-KM (variable-selection heuristic for K-means clustering) procedure (Brusco and Cradit, 2001). To identify outliers, we used a hybrid approach combining a clustering based approach and distance based approach. Simulation results indicate that the proposed automated K-means clustering procedure is effective to select variables and identify outliers. The implemented R program can be obtained at http://www.knou.ac.kr/~sskim/SVOKmeans.r.

로버스트 추정법을 이용한 자기상관회귀모형에서의 특이치 검출 (Outlier Detection of Autoregressive Models Using Robust Regression Estimators)

  • 이동희;박유성;김기환
    • 응용통계연구
    • /
    • 제19권2호
    • /
    • pp.305-317
    • /
    • 2006
  • 시계열 자료에서의 특이치, 특히 이 가운데 가법적 특이치가 모형의 식별, 모수의 추정 및 예측과 관련된 분석 전과정을 왜곡하는 것은 잘 알려져 있다. 그러나 특이치가 다수 발생하는 경우, 특히 연속적으로 집단을 이루어 발생할 때 대부분 특이치 검출방법은 가면화효과와 수렁화효과때문에 이들을 정확히 판별하지 못한다. 본 논문에서는 p차 자기상관회귀모형에 대한 고붕괴점 회귀추정량을 이용한 양방향 로버스트 필터방법을 제안했다. 실제 사례와 모의실험을 통해 제안한 방법이 매우 정확하게 시계열 자료에 포함된 특이치들을 검출하고 있음을 확인할 수 있다.

Diagnosis of Observations after Fit of Multivariate Skew t-Distribution: Identification of Outliers and Edge Observations from Asymmetric Data

  • Kim, Seung-Gu
    • 응용통계연구
    • /
    • 제25권6호
    • /
    • pp.1019-1026
    • /
    • 2012
  • This paper presents a method for the identification of "edge observations" located on a boundary area constructed by a truncation variable as well as for the identification of outliers and the after fit of multivariate skew $t$-distribution(MST) to asymmetric data. The detection of edge observation is important in data analysis because it provides information on a certain critical area in observation space. The proposed method is applied to an Australian Institute of Sport(AIS) dataset that is well known for asymmetry in data space.

Outlier Detection Based on Discrete Wavelet Transform with Application to Saudi Stock Market Closed Price Series

  • RASHEDI, Khudhayr A.;ISMAIL, Mohd T.;WADI, S. Al;SERROUKH, Abdeslam
    • The Journal of Asian Finance, Economics and Business
    • /
    • 제7권12호
    • /
    • pp.1-10
    • /
    • 2020
  • This study investigates the problem of outlier detection based on discrete wavelet transform in the context of time series data where the identification and treatment of outliers constitute an important component. An outlier is defined as a data point that deviates so much from the rest of observations within a data sample. In this work we focus on the application of the traditional method suggested by Tukey (1977) for detecting outliers in the closed price series of the Saudi Arabia stock market (Tadawul) between Oct. 2011 and Dec. 2019. The method is applied to the details obtained from the MODWT (Maximal-Overlap Discrete Wavelet Transform) of the original series. The result show that the suggested methodology was successful in detecting all of the outliers in the series. The findings of this study suggest that we can model and forecast the volatility of returns from the reconstructed series without outliers using GARCH models. The estimated GARCH volatility model was compared to other asymmetric GARCH models using standard forecast error metrics. It is found that the performance of the standard GARCH model were as good as that of the gjrGARCH model over the out-of-sample forecasts for returns among other GARCH specifications.

정상 시계열에서의 이상치 발견과 시계열 모형구축 (Outlier detection and time series modelling in the stationary time series)

  • 이종협;최기헌
    • 응용통계연구
    • /
    • 제5권2호
    • /
    • pp.139-156
    • /
    • 1992
  • 최근에 시계열에서의 이상치 발견을 위한 여러 가지 반복적인 방법들이 소개되었으나 이들 대부분은 시계열의 기저모형이 알려져 있거나 식별될 수 있다는 가정하에서 개발되었다. 그 렇지만 실제로 이상치들이 모형식별을 왜곡 시키거나 심지어는 불가능하게 만드는 경우가 발생한다. 본 논문에서는 두 개의 시계열 관측치 사이의 거리에 근거한 새로운 척도를 이용 한 이상치 탐색 방법을 제시하였다. 특히 이방법은 이상치를 발견하는데 시계열 모형에 의 존하지 않는다. 제안된 통계량에 대한 여러 가지 성질을 밝혔으며 이상치의 형태를 구별하 기 위해 전이함수모형을 이용하였다. 그밖에 이상치를 포함하고 있는 시계열의 모형을 구축 하기 위한 반복적인 절차를 제안했다.

  • PDF

Skew Normal Boxplot and Outliers

  • Huh, Myung-Hoe;Lee, Yong-Goo
    • Communications for Statistical Applications and Methods
    • /
    • 제19권4호
    • /
    • pp.591-595
    • /
    • 2012
  • We frequently use Tukey's boxplot to identify outliers in the batch of observations of the continuous variable. In doing so, we implicitly assume that the underlying distribution belongs to the family of normal distributions. Such a practice of data handling is often superficial and improper, since in reality too many variables manifest the skewness. In this short paper, we build a modified boxplot and set the outlier identification procedure by assuming that the observations are generated from the skew normal distribution (Azzalini, 1985), which is an extension of the normal distribution. Statistical performance of the proposed procedure is examined with simulated datasets.