• Title/Summary/Keyword: C4.5 Decision Tree

Search Result 84, Processing Time 0.04 seconds

A study on analysis of factors on in-hospital mortality for community-acquired pneumonia (지역사회획득 폐렴 환자의 퇴원시 사망 요인 분석)

  • Kim, Yoo-Mi
    • Journal of the Korean Data and Information Science Society
    • /
    • v.22 no.3
    • /
    • pp.389-400
    • /
    • 2011
  • This study was carried out to analysis factors related to in-hospital mortality of community-acquired peumonia using administrative database. The subjects were 5,353 community-acquired pneumonia inpatients of the Korean National Hospital Discharge Injury Survey 2004-2006 data. The data were analyzed using chi-squared test and decision tree model in the data mining technique. Among the decision tree model, C4.5 had the best performance. The critical factors on in-hospital mortality of communityacquired pneumonia are admission route, respiratory failure, congenital heart failure including age, comorbidity, and bed size. This study was carried out using the administrative database including patients' characteristics and comorbidity. However further study should be extensively including hospital characteristics, regional medical resources, and patient management practice behavior.

A Study on Split Variable Selection Using Transformation of Variables in Decision Trees

  • Chung, Sung-S.;Lee, Ki-H.;Lee, Seung-S.
    • Journal of the Korean Data and Information Science Society
    • /
    • v.16 no.2
    • /
    • pp.195-205
    • /
    • 2005
  • In decision tree analysis, C4.5 and CART algorithm have some problems of computational complexity and bias on variable selection. But QUEST algorithm solves these problems by dividing the step of variable selection and split point selection. When input variables are continuous, QUEST algorithm uses ANOVA F-test under the assumption of normality and homogeneity of variances. In this paper, we investigate the influence of violation of normality assumption and effect of the transformation of variables in the QUEST algorithm. In the simulation study, we obtained the empirical powers of variable selection and the empirical bias of variable selection after transformation of variables having various type of underlying distributions.

  • PDF

A Comparative Study on The Effective Use of Decision Tree Algorithms (의사결정 트리의 효용성 제고 방안에 관한 비교 연구)

  • Sug, Hyon-Tai
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2009.01a
    • /
    • pp.321-324
    • /
    • 2009
  • 비교적 적은 크기이면서 예측력에 있어 만족할 만한 의사결정목을 생성하는 방법으로서 적절한 크기의 샘플링을 제안하였다. 일반적으로 샘플의 크기가 작을수록 작은 의사결정목이 생성되므로 적절한 예측 정확도를 갖는 작은 트리를 생성하기를 원할 경우 적당한 크기의 샘플링을 하는 것이 트리의 최적화를 위한 계산을 더 시행하는 것보다 바람직하다고 할 수 있으며, 이와 같은 사실은 현재 알려진 가장 대표적 의사결정목 생성 알고리즘인 C4.5 및 CART를 사용하여 실험으로서 보여주었다.

  • PDF

A Study on the Design of Tolerance for Process Parameter using Decision Tree and Loss Function (의사결정나무와 손실함수를 이용한 공정파라미터 허용차 설계에 관한 연구)

  • Kim, Yong-Jun;Chung, Young-Bae
    • Journal of Korean Society of Industrial and Systems Engineering
    • /
    • v.39 no.1
    • /
    • pp.123-129
    • /
    • 2016
  • In the manufacturing industry fields, thousands of quality characteristics are measured in a day because the systems of process have been automated through the development of computer and improvement of techniques. Also, the process has been monitored in database in real time. Particularly, the data in the design step of the process have contributed to the product that customers have required through getting useful information from the data and reflecting them to the design of product. In this study, first, characteristics and variables affecting to them in the data of the design step of the process were analyzed by decision tree to find out the relation between explanatory and target variables. Second, the tolerance of continuous variables influencing on the target variable primarily was shown by the application of algorithm of decision tree, C4.5. Finally, the target variable, loss, was calculated by a loss function of Taguchi and analyzed. In this paper, the general method that the value of continuous explanatory variables has been used intactly not to be transformed to the discrete value and new method that the value of continuous explanatory variables was divided into 3 categories were compared. As a result, first, the tolerance obtained from the new method was more effective in decreasing the target variable, loss, than general method. In addition, the tolerance levels for the continuous explanatory variables to be chosen of the major variables were calculated. In further research, a systematic method using decision tree of data mining needs to be developed in order to categorize continuous variables under various scenarios of loss function.

Analysis of Leaf Node Ranking Methods for Spatial Event Prediction (의사결정트리에서 공간사건 예측을 위한 리프노드 등급 결정 방법 분석)

  • Yeon, Young-Kwang
    • Journal of the Korean Association of Geographic Information Studies
    • /
    • v.17 no.4
    • /
    • pp.101-111
    • /
    • 2014
  • Spatial events are predictable using data mining classification algorithms. Decision trees have been used as one of representative classification algorithms. And they were normally used in the classification tasks that have label class values. However since using rule ranking methods, spatial prediction have been applied in the spatial prediction problems. This paper compared rule ranking methods for the spatial prediction application using a decision tree. For the comparison experiment, C4.5 decision tree algorithm, and rule ranking methods such as Laplace, M-estimate and m-branch were implemented. As a spatial prediction case study, landslide which is one of representative spatial event occurs in the natural environment was applied. Among the rule ranking methods, in the results of accuracy evaluation, m-branch showed the better accuracy than other methods. However in case of m-brach and M-estimate required additional time-consuming procedure for searching optimal parameter values. Thus according to the application areas, the methods can be selectively used. The spatial prediction using a decision tree can be used not only for spatial predictions, but also for causal analysis in the specific event occurrence location.

Pridict of Liver cirrhosis susceptibility using Decision tree with SNP (Decision Tree와 SNP정보를 이용한 간경화 환자의 감수성 예측)

  • Kim, Dong-Hoi;Uhmn, Saang-Yong;Cho, Sung-Won;Ham, Ki-Baek;Kim, Jin
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2006.10a
    • /
    • pp.63-66
    • /
    • 2006
  • 본 논문에서는 SNP데이터를 이용하여 간경화에 대한 감수성을 예측하기 위해 의사결정 트리를 이용하였다. 데이터는 간경화 환자와 정상환자 총 116명의 데이터를 사용하였으며, Feature 값으로는 간질환과 밀접한 연관성을 갖는 28개의 SNP데이터를 사용하였다. 실험방법은 각각의 SNP에 대하여 의사결정트리로 분류율을 측정한 후 가장 높은 분류율을 가지는 SNP부터 조합해 나가는 방식으로 C4.5 의사결정트리를 이용 leave-one-out cross validation으로 간경화와 정상을 구분하는 정확도를 측정하였다. 실험결과 간 질환 관련 SNP중 IL1RN-S130S, IRNGR2-Q64R, IL-10(-592), IL1B_S35S 4개의 SNP조합에서 65.52%의 정확도를 얻을 수 있었다.

  • PDF

A data extension technique to handle incomplete data (불완전한 데이터를 처리하기 위한 데이터 확장기법)

  • Lee, Jong Chan
    • Journal of the Korea Convergence Society
    • /
    • v.12 no.2
    • /
    • pp.7-13
    • /
    • 2021
  • This paper introduces an algorithm that compensates for missing values after converting them into a format that can represent the probability for incomplete data including missing values in training data. In the previous method using this data conversion, incomplete data was processed by allocating missing values with an equal probability that missing variables can have. This method applied to many problems and obtained good results, but it was pointed out that there is a loss of information in that all information remaining in the missing variable is ignored and a new value is assigned. On the other hand, in the new proposed method, only complete information not including missing values is input into the well-known classification algorithm (C4.5), and the decision tree is constructed during learning. Then, the probability of the missing value is obtained from this decision tree and assigned as an estimated value of the missing variable. That is, some lost information is recovered using a lot of information that has not been lost from incomplete learning data.

Analysis of Factors for Seasonal Meat Color Characteristics in Hanwoo(Korean Cattle) Beef using Decision Tree Method (의사결정나무분석기법을 이용한 계절별 한우육의 육색 특성에 미치는 요인분석)

  • Kim, Seok-Jung;Kim, Yong-Sun;Song, Young-Han;Lee, Sung-Ki
    • Journal of Animal Science and Technology
    • /
    • v.44 no.5
    • /
    • pp.607-616
    • /
    • 2002
  • This study analyzed the effects of pH, sex, backfat thickness, ribeye area, cold carcass weight, shipping month, muscle internal temperature, average daily temperature, and average relative humidity for slaughtered Hanwoo to meat color by season. The analyses focused on interaction and each effect to meat color of the factors. For the result for analysis of multiple linear regressions, meat color values were decreased as pH increased in all meat color, and the meat color values increased as the backfat thickness was increased. As the results of the decision tree analysis by each factor, cow and steer slaughtered in spring and autumn were the highest in the lightness(L*). The redness(a*) was the cases that pH was less than 5.63 and average relative humidity was over than 71.5% for Hanwoo slaughtered in autumn. The chroma(C*) value was the highest for Hanwoo that was slaughtered in summer and autumn, the pH was less than 5.60, and the back fat thickness was over than 8 mm. The hue angle($h^0$) was shown that the muscle internal temperature was less than 4.7$^{\circ}C$ among Hanwoo which was slaughtered in spring, summer, and autumn, the pH was less than 5.66, and the back fat thickness was over than 8 mm.

Development of Decision Tree Software and Protein Profiling using Surface Enhanced laser Desorption/lonization - Time of Flight - Mass Spectrometry (SELDI-TOF-MS) in Papillary Thyroid Cancer (의사결정트리 프로그램 개발 및 갑상선유두암에서 질량분석법을 이용한 단백질 패턴 분석)

  • Yoon, Joon-Kee;Lee, Jun;An, Young-Sil;Park, Bok-Nam;Yoon, Seok-Nam
    • Nuclear Medicine and Molecular Imaging
    • /
    • v.41 no.4
    • /
    • pp.299-308
    • /
    • 2007
  • Purpose: The aim of this study was to develop a bioinformatics software and to test it in serum samples of papillary thyroid cancer using mass spectrometry (SELDI-TOF-MS). Materials and Methods: Development of 'Protein analysis' software performing decision tree analysis was done by customizing C4.5. Sixty-one serum samples from 27 papillary thyroid cancer, 17 autoimmune thyroiditis, 17 controls were applied to 2 types of protein chips, CM10 (weak cation exchange) and IMAC3 (metal binding - Cu). Mass spectrometry was performed to reveal the protein expression profiles. Decision trees were generated using 'Protein analysis' software, and automatically detected biomarker candidates. Validation analysis was performed for CM10 chip by random sampling. Results: Decision tree software, which can perform training and validation from profiling data, was developed. For CM10 and IMAC3 chips, 23 of 113 and 8 of 41 protein peaks were significantly different among 3 groups (p<0.05), respectively. Decision tree correctly classified 3 groups with an error rate of 3.3% for CM10 and 2.0% for IMAC3, and 4 and 7 biomarker candidates were detected respectively. In 2 group comparisons, all cancer samples were correctly discriminated from non-cancer samples (error rate = 0%) for CM10 by single node and for IMAC3 by multiple nodes. Validation results from 5 test sets revealed SELDI-TOF-MS and decision tree correctly differentiated cancers from non-cancers (54/55, 98%), while predictability was moderate in 3 group classification (36/55, 65%). Conclusion: Our in-house software was able to successfully build decision trees and detect biomarker candidates, therefore it could be useful for biomarker discovery and clinical follow up of papillary thyroid cancer.

Intelligent Service Reasoning Model Using Data Mining In Smart Home Environments (스마트 홈 환경에서 데이터 마이닝 기법을 이용한 지능형 서비스 추론 모델)

  • Kang, Myung-Seok;Kim, Hag-Bae
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.32 no.12B
    • /
    • pp.767-778
    • /
    • 2007
  • In this paper, we propose a Intelligent Service Reasoning (ISR) model using data mining in smart home environments. Our model creates a service tree used for service reasoning on the basis of C4.5 algorithm, one of decision tree algorithms, and reasons service that will be offered to users through quantitative weight estimation algorithm that uses quantitative characteristic rule and quantitative discriminant rule. The effectiveness in the performance of the developed model is validated through a smart home-network simulation.