• 제목/요약/키워드: Outliers test

검색결과 114건 처리시간 0.028초

군집 알고리즘을 이용한 순차적 이상치 탐지법 (A sequential outlier detecting method using a clustering algorithm)

  • 서한손;윤민
    • 응용통계연구
    • /
    • 제29권4호
    • /
    • pp.699-706
    • /
    • 2016
  • 검정절차가 생략된 이상치 탐지법은 구조적으로 수렁효과나 가면효과에 취약하기 때문에 다수의 이상치를 제대로 탐지하지 못할 때가 있다. 본 연구에서는 군집화에 의하여 구분된 소수 관찰치군을 이상치로 판정하는 방법에 보완될 검정절차를 다룬다. 이에 관련된 일반적인 방법은 탐지된 이상치 후보군의 개별적인 관찰치에 대해 다양한 종류의 t-검정을 수행하는 것이다. 본 연구에서는 이상치 후보군에 대한 검정을 수행하고 군집나무의 절단기준을 변경시켜 새로운 이상치군을 탐색해 나가는 순차적인 방법을 제안한다. 예제와 모의실험을 통해 제시된 방법과 기존의 방법들을 비교한다.

Low Outliers를 고려한 홍수빈도분석에 관한 연구 (A study on the Flood Frequency Analyzed in Consideration of Low Outliers.)

  • 이순혁;홍성표;박명근
    • 한국농공학회지
    • /
    • 제30권4호
    • /
    • pp.62-70
    • /
    • 1988
  • This study was conducted to solve the problems for the unsuitable parameters and the uncertainty of design flood can be appeared by low outliers were inclined to the lower part from the trend of the balance of the data. Derivation of reasonable design flood was attempted finally by modification of low outliers with analysis of flood frequency by means of Log Pearson Type Ill distribution. Three subwatersheds were selected as studying basins with the annual maximum series including low outliers along Geum River basin. The results through this study were analyzed and summarized as follows. 1. Log Pearson Type In distribution was confirmed as a reasonable one by X$^2$ goodness of fit test at Gong Ju, Gyu Am, og Cheon watershed along Geum River basin. 2. Probable flood flows for each watershed were derivated by flood frequency curve with outliers. 3. Weighted skew coefficient for each watershed was calculated for the evaluation of freq- uency factor which is needed for the modification of low outlier. 4. It was confirrned that adjusted frequency curve has a lower tendency than that of deletion of low outlier in common at all watersheds. 5. Final probable flood flows were derivated by modification with evaluation of modified basic statistics for three watersheds. 6. In comparison with a frequency curve with modification and one with outlier, The former has a higher probable flood flow within three years of return periods than that of the latter, and vice versa over three years of return periods.

  • PDF

Robust CUSUM test for time series of counts and its application to analyzing the polio incidence data

  • Kang, Jiwon
    • Journal of the Korean Data and Information Science Society
    • /
    • 제26권6호
    • /
    • pp.1565-1572
    • /
    • 2015
  • In this paper, we analyze the polio incidence data based on the Poisson autoregressive models, focusing particularly on change-point detection. Since the data include some strongly deviating observations, we employ the robust cumulative sum (CUSUM) test proposed by Kang and Song (2015) to perform the test for parameter change. Contrary to the result of Kang and Lee (2014), our data analysis indicates that there is no significant change in the case of the CUSUM test with strong robustness and the same result is obtained after ridding the polio data of outliers. We additionally consider the comparison of the forecasting performance. All the results demonstrate that the robust CUSUM test performs adequately in the presence of seemingly outliers.

치의학 연구에서 이상치의 처리 (Outlier detection in dental research)

  • 김기열
    • 대한치과의사협회지
    • /
    • 제55권9호
    • /
    • pp.604-616
    • /
    • 2017
  • In clinical dental research, errors occur in spite of careful study design and conduct. Data cleaning procedures intend to identify and correct these errors or at least to minimize their influence on study. Outlier is the one of these errors. Outlier detection is the first step in data analysis process which has a serious effect in the field of dental research. Hence, this paper aims to introduce the methods to detect the outliers and to examine their influences in statistical data analysis.

  • PDF

RAM 분석 정확도 향상을 위한 야전운용 데이터의 이상값과 결측값 처리 방안 (Method of Processing the Outliers and Missing Values of Field Data to Improve RAM Analysis Accuracy)

  • 김인석;정원
    • 한국신뢰성학회지:신뢰성응용연구
    • /
    • 제17권3호
    • /
    • pp.264-271
    • /
    • 2017
  • Purpose: Field operation data contains missing values or outliers due to various causes of the data collection process, so caution is required when utilizing RAM analysis results by field operation data. The purpose of this study is to present a method to minimize the RAM analysis error of the field data to improve the accuracy. Methods: Statistical methods are presented for processing of the outliers and the missing values of the field operating data, and after analyzing the RAM, the differences between before and after applying the technique are discussed. Results: The availability is estimated to be lower by 6.8 to 23.5% than that before processing, and it is judged that the processing of the missing values and outliers greatly affect the RAM analysis result. Conclusion: RAM analysis of OO weapon system was performed and suggestions for improvement of RAM analysis were presented through comparison with the new and current method. Data analysis results without appropriate treatment of error values may result in incorrect conclusions leading to inappropriate decisions and actions.

선형모형에서 특정 이상치 후보군에 대한 검정 (A Test on a Specific Set of Outlier Candidates in a Linear Model)

  • 서한손;윤민
    • 응용통계연구
    • /
    • 제27권2호
    • /
    • pp.307-315
    • /
    • 2014
  • 이상치 후보군을 검정할 때 일반적으로 정확한 검정 통계량의 분포가 존재하지 않는다. 이에 따라 전체 관찰치군에 대한 검정대신 개별 관찰치에 대한 검정을 수행하거나 실험에 의해 계산된 유의값을 사용하여 이상치 가설검정을 수행한다. 본 연구에서는 임의의 관찰치 집단 또는 이상치 탐지절차에 따라 이상치 후보로 탐지된 특정 관찰치 집단의 이상치 여부를 검정하는 방법을 제시한다. 제시된 방법은 기존의 이상치 탐지기법에서 사용되는 검정방법과 모의실험을 통해 검정력을 비교한다.

Outlier Tests in Sample Surveys

  • Namkyung, Pyong;Lee, Joon Suk
    • Communications for Statistical Applications and Methods
    • /
    • 제7권2호
    • /
    • pp.447-456
    • /
    • 2000
  • In this paper, we considered three methods for outlier identification sample surveys. First, we studied method of handling and adjusting outliers in normal population. Second, we studied existing methods using mean, maximum and minimum and proposed a test using of median which well reflects characteristic of data regardless of sampling distribution. Finally, we showed our test using median works better than Dixon and mean test through simulation.

  • PDF

Test for Parameter Change based on the Estimator Minimizing Density-based Divergence Measures

  • Na, Ok-Young;Lee, Sang-Yeol;Park, Si-Yun
    • 한국통계학회:학술대회논문집
    • /
    • 한국통계학회 2003년도 춘계 학술발표회 논문집
    • /
    • pp.287-293
    • /
    • 2003
  • In this paper we consider the problem of parameter change based on the cusum test proposed by Lee et al. (2003). The cusum test statistic is constructed utilizing the estimator minimizing density-based divergence measures. It is shown that under regularity conditions, the test statistic has the limiting distribution of the sup of standard Brownian bridge. Simulation results demonstrate that the cusum test is robust when there arc outliers.

  • PDF

On a Robust Test for Parallelism of Regression Lines against Ordered Alternatives

  • Song, Moon-Sup;Kim, Jin-Ho
    • Communications for Statistical Applications and Methods
    • /
    • 제4권2호
    • /
    • pp.565-579
    • /
    • 1997
  • A robust test is proposed for the problem of testing the parallelism of several regression lines against ordered alternatives. The proposed test statistic is based on a linear combination of one-step pairwise GM-estimators. We compare the performance of the proposed test with that of the other tests through a Monte Carlo simulation. The results of the simulation study show that the proposed test has stable levels, good empirical powers in various circumstances, and particularly higher empirical powers under the presence of extreme outliers or leverage points.

  • PDF

평균이동모형을 이용한 성장곡선모형의 이상점 진단에 관한 연구 (Outlier Detection in Growth Curve Model Using Mean-Shift Model)

  • 심규박
    • Journal of the Korean Data and Information Science Society
    • /
    • 제10권2호
    • /
    • pp.369-385
    • /
    • 1999
  • 성장곡선모형에서 다중 이상값들이나 영향관측값들을 탐지하는 문제는 선형회귀모형에서의 문제에 비해 매우 복잡하여 거의 이루어지지 않고 있는 실정이다. 본 연구에서는 이상점을 포함하고 있는 성장곡선모형에서 이들을 탐지하는 방법으로 평균이동모형을 이용하는 방법을 소개하였다. 이 방법을 이용하여 찾아낸 자료가 이상점인지의 여부를 예측표본재이용 의사 베이즈 우도 기준법을 이용한 등분산성의 검정을 통해 알아보았다. 끝으로 Potthoff(1964)등이 사용한 자료를 이용한 예제를 통해 이상점 탐지와 등분 산성 검정을 실시한 결과를 제시하였다.

  • PDF