• 제목/요약/키워드: Classification and Regression tree

검색결과 211건 처리시간 0.025초

단순작업으로 인한 정신피로도 측정을 위한 음성기술을 이용한 CART 기반 진단모델 (A CART-based diagnostic model using speech technology for evaluating mental fatigue caused by monotonous work)

  • 권철홍
    • 말소리와 음성과학
    • /
    • 제8권4호
    • /
    • pp.97-101
    • /
    • 2016
  • This paper presents a CART(Classification and Regression Tree)-based model to diagnose mental fatigue using speech technology. The parameters used in the model are the significant speech parameters highly correlated to the fatigue and questionnaire responses obtained before and after imposing the fatigue. It is shown from the experiments that the proposed model achieves classification accuracies of 96.67% and 98.33% using the speech parameters and questionnaire responses, respectively. This implies that the proposed model can be used as a tool to diagnose the mental fatigue, and that speech technology is useful to diagnose the fatigue.

기계학습기법을 이용한 땅밀림 위험등급 분류 (Classification of Soil Creep Hazard Class Using Machine Learning)

  • 이기하;레수안히엔;연민호;서준표;이창우
    • 한국방재안전학회논문집
    • /
    • 제14권3호
    • /
    • pp.17-27
    • /
    • 2021
  • 본 연구에서는 6개의 기계학습 기법들을 활용하여 2019년과 2020년 전국 땅밀림 현장조사 결과를 기반으로 땅밀림 위험지역을 A부터 C까지 3개 등급(A등급: 위험, B등급: 보통, C등급: 양호)으로 구분할 수 있는 분류모형을 구축하고, 분류 정확도를 비교·분석한다. 기계학습 기법으로는 K-Nearest Neighbor, Support Vector Machine, Logistic Regression, Decision Tree, Random Forest, Extreme Gradient Boosting 총 6개를 적용하였다. 분류 정확도 분석결과, 6개의 기법 모두 0.9 이상의 우수한 정확도를 보여주었다. 수치형 자료를 학습에 적용한 경우가, 문자형 자료를 학습한 모형보다 우수한 성능을 나타냈으며, 현장조사 평가점수 자료군(C1~C4) 보다는 전문가의견이 반영된 평가점수 자료군(R1~R4)으로 학습한 모형이 정확도가 높은 것으로 분석되었다. 특히, 직접징후와 간접징후 정보를 학습에 반영한 경우가 예측정확도가 높게 나타났다. 향후 땅밀림 현장조사 자료가 지속적으로 확보될 경우, 본 연구에서 활용한 기계학습기법은 땅밀림 분류를 위한 도구로 활용이 가능할 것으로 판단된다.

의사결정나무 분류와 인공신경망을 이용한 토양수분 산정모형 개발 (Development of a Soil Moisture Estimation Model Using Artificial Neural Networks and Classification and Regression Tree(CART))

  • 김광섭;박정아
    • 대한토목학회논문집
    • /
    • 제31권2B호
    • /
    • pp.155-163
    • /
    • 2011
  • 본 연구에서는 의사결정나무(CART)기법, 인공신경망모형, 인공위성 원격탐사자료와 지형자료 및 지상 기상관측망자료를 이용하여 토양수분을 산정하는 모형을 개발하였다. 본 모형의 검증을 위하여 사용된 토양수분 관측자료는 용담댐 유역에서 관측된 5개 지점의 토양수분자료를 사용하였다. 가용자료에 대해 CART기법을 적용하여 자료를 분류한 다음 분류된 각 자료집단에 대하여 인공신경망(Artificial Neural Networks)모형을 적용하여 토양수분 분포를 예측하였다. 모형의 학습에 사용된 주천, 부귀, 상전, 안천 지점의 토양수분 산정치는 관측치와 약 0.92-0.96의 상관계수, 약 1.00-1.88%의 평균제곱근오차와 약 0.75-1.45%의 평균절대오차를 보여주었다. 토양수분 추정모형을 검증하기 위해 천천2의 지점에 적용한 결과 약 0.91의 상관계수, 약 3.19%의 평균제곱근오차, 약 2.72%의 평균절대오차를 보여 CART기법과 인공신경망모형을 연계한 토양수분 산정모형이 토양수분 분포제시 활용에 적절한 것으로 판단된다.

머신러닝을 통한 잉크 필요량 예측 알고리즘 (Machine Learning Algorithm for Estimating Ink Usage)

  • 권세욱;현영주;태현철
    • 산업경영시스템학회지
    • /
    • 제46권1호
    • /
    • pp.23-31
    • /
    • 2023
  • Research and interest in sustainable printing are increasing in the packaging printing industry. Currently, predicting the amount of ink required for each work is based on the experience and intuition of field workers. Suppose the amount of ink produced is more than necessary. In this case, the rest of the ink cannot be reused and is discarded, adversely affecting the company's productivity and environment. Nowadays, machine learning models can be used to figure out this problem. This study compares the ink usage prediction machine learning models. A simple linear regression model, Multiple Regression Analysis, cannot reflect the nonlinear relationship between the variables required for packaging printing, so there is a limit to accurately predicting the amount of ink needed. This study has established various prediction models which are based on CART (Classification and Regression Tree), such as Decision Tree, Random Forest, Gradient Boosting Machine, and XGBoost. The accuracy of the models is determined by the K-fold cross-validation. Error metrics such as root mean squared error, mean absolute error, and R-squared are employed to evaluate estimation models' correctness. Among these models, XGBoost model has the highest prediction accuracy and can reduce 2134 (g) of wasted ink for each work. Thus, this study motivates machine learning's potential to help advance productivity and protect the environment.

Practice of causal inference with the propensity of being zero or one: assessing the effect of arbitrary cutoffs of propensity scores

  • Kang, Joseph;Chan, Wendy;Kim, Mi-Ok;Steiner, Peter M.
    • Communications for Statistical Applications and Methods
    • /
    • 제23권1호
    • /
    • pp.1-20
    • /
    • 2016
  • Causal inference methodologies have been developed for the past decade to estimate the unconfounded effect of an exposure under several key assumptions. These assumptions include, but are not limited to, the stable unit treatment value assumption, the strong ignorability of treatment assignment assumption, and the assumption that propensity scores be bounded away from zero and one (the positivity assumption). Of these assumptions, the first two have received much attention in the literature. Yet the positivity assumption has been recently discussed in only a few papers. Propensity scores of zero or one are indicative of deterministic exposure so that causal effects cannot be defined for these subjects. Therefore, these subjects need to be removed because no comparable comparison groups can be found for such subjects. In this paper, using currently available causal inference methods, we evaluate the effect of arbitrary cutoffs in the distribution of propensity scores and the impact of those decisions on bias and efficiency. We propose a tree-based method that performs well in terms of bias reduction when the definition of positivity is based on a single confounder. This tree-based method can be easily implemented using the statistical software program, R. R code for the studies is available online.

의사결정나무분석법을 활용한 6차산업 유형별 산업적 기능결합 요인탐색 (Exploring Industrial Function Combining Factors for Each Type in the 6th Industry Based on Decision Tree Analysis)

  • 김정태
    • 농촌지도와개발
    • /
    • 제23권3호
    • /
    • pp.243-255
    • /
    • 2016
  • This study aims to identify the characteristics of businesses influencing the choice of their type in the 6th industry and analyze how they work. This study analyzed data of 752 businesses certified as belonging to the 6th industry in 2015 through the classification and regression tree (CART) algorithm in decision tree analysis. The results of analysis showed that the type of agricultural product processing, region, the type of service, and the production percentage in a province affected a choice of the type. The most important variable that impacted how businesses in the 6th industry chose their type was the type of agricultural product processing, and if a business produced simple agricultural products, it was likely to specialize into $1st^*2nd$ or $1st^*3rd$. Access to large consumption areas was a critical factor in the growth of 2nd and 3rd industrial functions. These findings would contribute to establishing a model to develop the 6th industry and empirically demonstrate the importance of access to large consumption areas for agricultural businesses and rural tourism.

온라인 무료 샘플 판촉의 효과적 활용을 위한 기계학습 기반 고객분류예측 모형 (A Machine Learning-based Customer Classification Model for Effective Online Free Sample Promotions)

  • 원하람;김무전;안현철
    • 한국정보시스템학회지:정보시스템연구
    • /
    • 제27권3호
    • /
    • pp.63-80
    • /
    • 2018
  • Purpose The purpose of this study is to build a machine learning-based customer classification model to promote customer expansion effect of the free sample promotion. Specifically, the proposed model classifies potential target customers who are expected to purchase the products included in the free sample promotion after receiving the free samples. Design/methodology/approach This study proposes to build a customer classification model for determining customers suitable for providing free samples by using various machine learning techniques such as logistic regression, multiple discriminant analysis, case-based reasoning, decision tree, artificial neural network, and support vector machine. To validate the usefulness of the proposed model, we apply it to a real-world free sample-based target marketing case of a Korean major cosmetic retail company. Findings Experimental results show that a machine learning-based customer classification model presents satisfactory accuracy ranging from 70% to 75%. In particular, support vector machine is found to be the most effective machine learning technique for free sample-based target marketing model. Our study sheds a light on customer relationship management strategies using free sample promotions.

데이터 마이닝을 이용한 시멘트 소성공정 질소산화물(NOx)배출 관리 방법에 관한 연구 (A Study on NOx Emission Control Methods in the Cement Firing Process Using Data Mining Techniques)

  • 박철홍;김용수
    • 품질경영학회지
    • /
    • 제46권3호
    • /
    • pp.739-752
    • /
    • 2018
  • Purpose: The purpose of this study was to investigate the relationship between kiln processing parameters and NOx emissions that occur in the sintering and calcination steps of the cement manufacturing process and to derive the main factors responsible for producing emissions outside emission limit criteria, as determined by category models and classification rules, using data mining techniques. The results from this study are expected to be useful as guidelines for NOx emission control standards. Methods: Data were collected from Precalciner Kiln No.3 used in one of the domestic cement plants in Korea. Thirty-four independent variables affecting NOx generation and dependent variables that exceeded or were below the NOx emiision limit (>1 and <0, respectively) were examined during kiln processing. These data were used to construct a detection model of NOx emission, in which emissions exceeded or were below the set limits. The model was validated using SPSS MODELER 18.0, artificial neural network, decision treee (C5.0), and logistic regression analysis data mining techniques. Results: The decision tree (C5.0) algorithm best represented NOx emission behavior and was used to identify 10 processing variables that resulted in NOx emissions outside limit criteria. Conclusion: The results of this study indicate that the decision tree (C5.0) can be applied for real-time monitoring and management of NOx emissions during the cement firing process to satisfy NOx emission control standards and to provide for a more eco-friendly cement product.

중규모급 단어 인식기의 실시간 구현을 위한 무감독 단어집단화 알고리듬 (Unsupervised Word Grouping Algorithm for real-time implementation of Medium vocabulary recognition)

  • 임동식;김진영;백성준
    • 한국음향학회:학술대회논문집
    • /
    • 한국음향학회 1999년도 학술발표대회 논문집 제18권 2호
    • /
    • pp.81-84
    • /
    • 1999
  • 본 논문에서는 중규모급 단어인식기의 실시간 구현을 위한 무감독 단어집단화 알고리듬을 제안한다. 무감독 단어집단화는 인식대상 어휘 수가 많은 대용량 음성인식 시스템에서 대상 어휘 수를 줄여주는 역할을 하는 전처리기의 성격을 갖는다. 무감독 집단화를 위해 각 단어의 유$\cdot$무성음 고유의 특성을 잘 반영할 수 있는 특징 파라미터 5개를 사용하여 패턴 인식과 회귀분석에서 널리 사용되고 있는 분류$\cdot$회귀트리(Classification And Regression Tree)에 적용시키는 방법으로 접근하였고, 각 단어의 frame 수를 일정하게 n개로 분할(segment)하여 1개의 tree를 생성시키는 방법과 각 segment에 해당하는 tree를 생성시켜 segment들 사이의 교집합 성분으로 단어들을 집단화 하였다 실험결과 탐색 대상단어 22개에서 평균2.21개로 줄어 전체 대상 단어의 $10\%$만을 탐색하여 인식할 수 있는 방법을 제시할 수 있었다.

  • PDF

변수선택 편향이 없는 회귀나무를 만들기 위한 알고리즘 (Regression Trees with. Unbiased Variable Selection)

  • 김진흠;김민호
    • 응용통계연구
    • /
    • 제17권3호
    • /
    • pp.459-473
    • /
    • 2004
  • 본 논문에서는 Breiman 등(1984)의 전체탐색법이 갖고 있는 변수선택 편향을 극복할 수 있는 알고리즘을 제안하였다. 제안한 알고리즘은 노드의 분리 변수를 선택하는 단계와 그 선택된 변수에 대해서만 이진분리를 위한 분리점을 찾는 단계로 나뉘어져 있다. 예측변수가 연속형 일 때는 스피어만의 순위상관계수에 의한 검정을 수행하고, 범주형일 때는 크루스칼-왈리스의 통계량에 의한 검정을 수행하여 통계적으로 가장 유의한 변수를 분리변수로 선택하였고 Breiman 등(1984)의 전체탐색법을 그 변수에만 적용하여 노드의 분리기준을 정하였다 모의실험 연구를 통해 Breiman등(19히)의 CART와 제안한 알고리즘을 변수선택 편의, 변수선택력파 평균제곱오차 측면에서 서로 비교하였다. 아울러 두 알고리즘을 실제 자료에 적용하여 효율을 서로 비교하였다.