• 제목/요약/키워드: Mixed Data Sampling

검색결과 109건 처리시간 0.019초

A Generalized Mixed-Effects Model for Vaccination Data

  • Choi, Jae-Sung
    • Journal of the Korean Data and Information Science Society
    • /
    • 제15권2호
    • /
    • pp.379-386
    • /
    • 2004
  • This paper deals with a mixed logit model for vaccination data. The effect of a newly developed vaccine for a certain chicken disease can be evaluated by a noninfection rate after injecting chicken with the disease vaccine. But there are a lot of factors that might affect the noninfecton rate. Some of these are fixed and others are random. Random factors are sometimes coming from the sampling scheme for choosing experimental units. This paper suggests a mixed model when some fixed factors need to have different experimental sizes by an experimental design and illustrates how to estimate parameters in a suggested model.

  • PDF

A Proportional Odds Mixed - Effects Model for Ordinal Data

  • Choi, Jae-Sung
    • Journal of the Korean Data and Information Science Society
    • /
    • 제18권2호
    • /
    • pp.471-479
    • /
    • 2007
  • This paper discusses about how to build up mixed-effects model for analysing ordinal response data by using cumulative logits. Random factors are assumed to be coming from the designed sampling scheme for choosing observational units. Since the observed responses of individuals are ordinal, a proportional odds model with two random effects is suggested. Estimation procedure for the unknown parameters in a suggested model is also discussed by an illustrated example.

  • PDF

불균형 데이터 처리를 통한 침입탐지 성능향상에 관한 연구 (A study on intrusion detection performance improvement through imbalanced data processing)

  • 정일옥;지재원;이규환;김묘정
    • 융합보안논문지
    • /
    • 제21권3호
    • /
    • pp.57-66
    • /
    • 2021
  • 침입탐지 분야에서 딥러닝과 머신러닝을 이용한 탐지성능이 검증되면서 이를 활용한 사례가 나날이 증가하고 있다. 하지만, 학습에 필요한 데이터 수집이 어렵고, 수집된 데이터의 불균형으로 인해 머신러닝 성능이 현실에 적용되는데 어려움이 있다. 본 논문에서는 이에 대한 해결책으로 불균형 데이터 처리를 위해 t-SNE 시각화를 이용한 혼합샘플링 기법을 제안한다. 이를 위해 먼저, 페이로드를 포함한 침입탐지 이벤트에 대해서 특성에 맞게 필드를 분리한다. 분리된 필드에 대해 TF-IDF 기반의 피처를 추출한다. 추출된 피처를 기반으로 혼합샘플링 기법을 적용 후 t-SNE를 이용한 데이터 시각화를 통해 불균형 데이터가 처리된 침입탐지에 최적화된 데이터셋을 얻게 된다. 공개 침입탐지 데이터셋 CSIC2012를 통해 9가지 샘플링 기법을 적용하였으며, 제안한 샘플링 기법이 F-score, G-mean 평가 지표를 통해 탐지성능이 향상됨을 검증하였다.

Subset 샘플링 검증 기법을 활용한 MSCRED 모델 기반 발전소 진동 데이터의 이상 진단 (Anomaly Detection In Real Power Plant Vibration Data by MSCRED Base Model Improved By Subset Sampling Validation)

  • 홍수웅;권장우
    • 융합정보논문지
    • /
    • 제12권1호
    • /
    • pp.31-38
    • /
    • 2022
  • 본 논문은 전문가 독립적 비지도 신경망 학습 기반 다변량 시계열 데이터 분석 모델인 MSCRED(Multi-Scale Convolutional Recurrent Encoder-Decoder)의 실제 현장에서의 적용과 Auto-encoder 기반인 MSCRED 모델의 한계인, 학습 데이터가 오염되지 않아야 된다는 점을 극복하기 위한 학습 데이터 샘플링 기법인 Subset Sampling Validation을 제시한다. 라벨 분류가 되어있는 발전소 장비의 진동 데이터를 이용하여 1) 학습 데이터에 비정상 데이터가 섞여 있는 상황을 재현하고, 이를 학습한 경우 2) 1과 같은 상황에서 Subset Sampling Validation 기법을 통해 학습 데이터에서 비정상 데이터를 제거한 경우의 Anomaly Score를 비교하여 MSCRED와 Subset Sampling Validation 기법을 유효성을 평가한다. 이를 통해 본 논문은 전문가 독립적이며 오류 데이터에 강한 이상 진단 프레임워크를 제시해, 다양한 다변량 시계열 데이터 분야에서의 간결하고 정확한 해결 방법을 제시한다.

Supremacy of Realized Variance MIDAS Regression in Volatility Forecasting of Mutual Funds: Empirical Evidence From Malaysia

  • WAN, Cheong Kin;CHOO, Wei Chong;HO, Jen Sim;ZHANG, Yuruixian
    • The Journal of Asian Finance, Economics and Business
    • /
    • 제9권7호
    • /
    • pp.1-15
    • /
    • 2022
  • Combining the strength of both Mixed Data Sampling (MIDAS) Regression and realized variance measures, this paper seeks to investigate two objectives: (1) evaluate the post-sample performance of the proposed weekly Realized Variance-MIDAS (RVar-MIDAS) in one-week ahead volatility forecasting against the established Generalized Autoregressive Conditional Heteroskedasticity (GARCH) model and the less explored but robust STES (Smooth Transition Exponential Smoothing) methods. (2) comparing forecast error performance between realized variance and squared residuals measures as a proxy for actual volatility. Data of seven private equity mutual fund indices (generated from 57 individual funds) from two different time periods (with and without financial crisis) are applied to 21 models. Robustness of the post-sample volatility forecasting of all models is validated by the Model Confidence Set (MCS) Procedures and revealed: (1) The weekly RVar-MIDAS model emerged as the best model, outperformed the robust DAILY-STES methods, and the weekly DAILY-GARCH models, particularly during a volatile period. (2) models with realized variance measured in estimation and as a proxy for actual volatility outperformed those using squared residual. This study contributes an empirical approach to one-week ahead volatility forecasting of mutual funds return, which is less explored in past literature on financial volatility forecasting compared to stocks volatility.

AN APPROACH TO THE TRAINING OF A SUPPORT VECTOR MACHINE (SVM) CLASSIFIER USING SMALL MIXED PIXELS

  • Yu, Byeong-Hyeok;Chi, Kwang-Hoon
    • 대한원격탐사학회:학술대회논문집
    • /
    • 대한원격탐사학회 2008년도 International Symposium on Remote Sensing
    • /
    • pp.386-389
    • /
    • 2008
  • It is important that the training stage of a supervised classification is designed to provide the spectral information. On the design of the training stage of a classification typically calls for the use of a large sample of randomly selected pure pixels in order to characterize the classes. Such guidance is generally made without regard to the specific nature of the application in-hand, including the classifier to be used. An approach to the training of a support vector machine (SVM) classifier that is the opposite of that generally promoted for training set design is suggested. This approach uses a small sample of mixed spectral responses drawn from purposefully selected locations (geographical boundaries) in training. A sample of such data should, however, be easier and cheaper to acquire than that suggested by traditional approaches. In this research, we evaluated them against traditional approaches with high-resolution satellite data. The results proved that it can be used small mixed pixels to derive a classification with similar accuracy using a large number of pure pixels. The approach can also reduce substantial costs in training data acquisition because the sampling locations used are commonly easy to observe.

  • PDF

Methods and Techniques for Variance Component Estimation in Animal Breeding - Review -

  • Lee, C.
    • Asian-Australasian Journal of Animal Sciences
    • /
    • 제13권3호
    • /
    • pp.413-422
    • /
    • 2000
  • In the class of models which include random effects, the variance component estimates are important to obtain accurate predictors and estimators. Variance component estimation is straightforward for balanced data but not for unbalanced data. Since orthogonality among factors is absent in unbalanced data, various methods for variance component estimation are available. REML estimation is the most widely used method in animal breeding because of its attractive statistical properties. Recently, Bayesian approach became feasible through Markov Chain Monte Carlo methods with increasingly powerful computers. Furthermore, advances in variance component estimation with complicated models such as generalized linear mixed models enabled animal breeders to analyze non-normal data.

소프트웨어 신뢰모형에 대한 베이지안 접근 (Bayesian Approach for Software Reliability Models)

  • 최기헌
    • Journal of the Korean Data and Information Science Society
    • /
    • 제10권1호
    • /
    • pp.119-133
    • /
    • 1999
  • 마코브체인 몬테칼로 방법을 소프트웨어 신뢰모형에 이용하였다. 베이지안 추론에서 조건부 분포를 가지고 사후분포를 결정하는데 있어서의 계산 문제를 고찰하였다. 특히 레코드값을 통계량을 갖고서 혼합과정과 중첩과정에 대하여 깁스샘플링 알고리즘과 메트로폴리스 알고리즘을 활용하여 베이지안 계산과 모형 선택을 제시하고 모의실험자료를 이용하여 수치적 인 계산을 시행하고 그 결과를 비교하였다.

  • PDF

Differential Mobility Analyzer(DMA)와 Condensation Nuclei Counter(CNC)를 이용한 입자크기 분포 측정에서 샘플링 튜브와 CNC에서의 혼합 효과가 입자 크기 분포 측정에 미치는 영향에 관한 연구 (Study on the Contribution of Mixing Effects in Sampling Tube and Condensation Nuclei Counter(CNC) to the measurement of size distribution obtained using Differential Mobility Analyzer and CNC)

  • 이윤수;안강호
    • 대한기계학회:학술대회논문집
    • /
    • 대한기계학회 2001년도 춘계학술대회논문집D
    • /
    • pp.104-109
    • /
    • 2001
  • The time to measure the size distribution using Condensation Nuclei Counter(CNC) and Differential Mobility Analyzer(DMA) can be shortened by classifying particles ramping the DMA voltage exponentially and continuously. In measurement, particles sampled at different time are mixed together going through sampling tube and CNC. Because the size distribution is inversed by using detector responses to sampling time intervals in this accelerated method, the mixing effects give inversion errors to the size distribution. The mixing effects can be considered by appling the transfer function with mixing effects to the data inversion. The inversion considering this effects gives birth to the size distribution shifted to the opposite direction of the size scanning.

  • PDF