• Title/Summary/Keyword: 변수선택 편의

Search Result 103, Processing Time 0.029 seconds

회귀나무에서 변수선택 편의에 관한 연구

  • Kim, Min-Ho;Kim, Jin-Heum
    • Proceedings of the Korean Statistical Society Conference
    • /
    • 2003.10a
    • /
    • pp.263-268
    • /
    • 2003
  • Breiman, Friedman, Olshen and Stone(1984)의 전체탐색법에 의한 회귀나무는 상대적으로 많은 분리가 가능한 변수로 분리기준이 정해지는 편의 현상을 갖고 있다. 본 연구에서는 이런 문제점을 해결할 수 있는 알고리즘을 제안하여 변수선택편의가 없는 회귀나무를 만들고자 한다. 제안하는 알고리즘은 노드의 분리변수를 선택하는 단계와 그 선택된 변수에 의해 이진분리를 위한 분리점을 찾는 단계로 구성되어 있다. 예측변수 중에서 목표변수와 가장 밀접하게 연관된 예측변수는 예측변수의 자료의 종류에 따라 스피어만의 순위상관계수에 의한 검정 혹은 크루스칼-왈리스의 통계량에 의한 검정을 수행하여 가장 통계적으로 유의한 변수로 선택하였고, 선택된 변수에만 Breiman et al.(1984)의 전체선택법을 적용하여 분리점을 결정하였다. 모의실험을 통해 변수선택편의, 변수선택력 , 그리고 평균제곱오차 측면에서 Breiman et al. (1984)의 CART(Classification and Regression Trees)와 제안한 알고리즘을 서로 비교하였다. 또한, 두 알고리즘을 실제 자료에 적용하여 효율을 서로 비교하였다.

  • PDF

A Study on Variable Selection Bias in Data Mining Software Packages (데이터마이닝 패키지에서 변수선택 편의에 관한 연구)

  • 송문섭;윤영주
    • The Korean Journal of Applied Statistics
    • /
    • v.14 no.2
    • /
    • pp.475-486
    • /
    • 2001
  • 데이터마이닝 패키지에 구현된 분류나무 알고리즘 가운데 CART, CHAID, QUEST, C4.5에서 변수 선택법을 비교하였다. CART의 전체탐색법이 편의를 갖는다는 사실은 잘알려졌으며, 여기서는 상품화된 패키지들에서 이들 알고리즘의 편의와 선택력을 모의실험 연구를 통하여 비교하였다. 상용 패키지로는 CART, Enterprise Miner, AnswerTree, Clementine을 사용하였다. 본 논문의 제한된 모의실험 연구 결과에 의하면 C4.5와 CART는 모두 변수선택에서 심각한 편의를 갖고 있으며, CHAID와 QUEST는 비교적 안정된 결과를 보여주고 있었다.

  • PDF

Regression Trees with. Unbiased Variable Selection (변수선택 편향이 없는 회귀나무를 만들기 위한 알고리즘)

  • 김진흠;김민호
    • The Korean Journal of Applied Statistics
    • /
    • v.17 no.3
    • /
    • pp.459-473
    • /
    • 2004
  • It has well known that an exhaustive search algorithm suggested by Breiman et. a1.(1984) has a trend to select the variable having relatively many possible splits as an splitting rule. We propose an algorithm to overcome this variable selection bias problem and then construct unbiased regression trees based on the algorithm. The proposed algorithm runs two steps of selecting a split variable and determining a split rule for binary split based on the split variable. Simulation studies were performed to compare the proposed algorithm with Breiman et a1.(1984)'s CART(Classification and Regression Tree) in terms of degree of variable selection bias, variable selection power, and MSE(Mean Squared Error). Also, we illustrate the proposed algorithm with real data sets.

A Study on Selection of Split Variable in Constructing Classification Tree (의사결정나무에서 분리 변수 선택에 관한 연구)

  • 정성석;김순영;임한필
    • The Korean Journal of Applied Statistics
    • /
    • v.17 no.2
    • /
    • pp.347-357
    • /
    • 2004
  • It is very important to select a split variable in constructing the classification tree. The efficiency of a classification tree algorithm can be evaluated by the variable selection bias and the variable selection power. The C4.5 has largely biased variable selection due to the influence of many distinct values in variable selection and the QUEST has low variable selection power when a continuous predictor variable doesn't deviate from normal distribution. In this thesis, we propose the SRT algorithm which overcomes the drawback of the C4.5 and the QUEST. Simulations were performed to compare the SRT with the C4.5 and the QUEST. As a result, the SRT is characterized with low biased variable selection and robust variable selection power.

A study on bias effect of LASSO regression for model selection criteria (모형 선택 기준들에 대한 LASSO 회귀 모형 편의의 영향 연구)

  • Yu, Donghyeon
    • The Korean Journal of Applied Statistics
    • /
    • v.29 no.4
    • /
    • pp.643-656
    • /
    • 2016
  • High dimensional data are frequently encountered in various fields where the number of variables is greater than the number of samples. It is usually necessary to select variables to estimate regression coefficients and avoid overfitting in high dimensional data. A penalized regression model simultaneously obtains variable selection and estimation of coefficients which makes them frequently used for high dimensional data. However, the penalized regression model also needs to select the optimal model by choosing a tuning parameter based on the model selection criterion. This study deals with the bias effect of LASSO regression for model selection criteria. We numerically describes the bias effect to the model selection criteria and apply the proposed correction to the identification of biomarkers for lung cancer based on gene expression data.

Joint penalization of components and predictors in mixture of regressions (혼합회귀모형에서 콤포넌트 및 설명변수에 대한 벌점함수의 적용)

  • Park, Chongsun;Mo, Eun Bi
    • The Korean Journal of Applied Statistics
    • /
    • v.32 no.2
    • /
    • pp.199-211
    • /
    • 2019
  • This paper is concerned with issues in the finite mixture of regression modeling as well as the simultaneous selection of the number of mixing components and relevant predictors. We propose a penalized likelihood method for both mixture components and regression coefficients that enable the simultaneous identification of significant variables and the determination of important mixture components in mixture of regression models. To avoid over-fitting and bias problems, we applied smoothly clipped absolute deviation (SCAD) penalties on the logarithm of component probabilities suggested by Huang et al. (Statistical Sinica, 27, 147-169, 2013) as well as several well-known penalty functions for coefficients in regression models. Simulation studies reveal that our method is satisfactory with well-known penalties such as SCAD, MCP, and adaptive lasso.

An Analysis of Job Selection, Major-Job Match and Wage Level of College Graduates (대학 졸업생의 직업선택과 임금 수준)

  • Park, Jae-Min
    • Journal of Korea Technology Innovation Society
    • /
    • v.14 no.1
    • /
    • pp.22-39
    • /
    • 2011
  • This study examines the wage level from a viewpoint of major-job match as part of an analysis on the skill mismatch problem in 4-year college graduates. The empirical analysis explicitly incorporate the sample selection bias as an econometric problem not only suggested but merely introduced in the earlier studies. This study also set up a major-job match variable, which was usually handled as a binary variable for analytical convenience, as a polychotomous choice variable in selection equation as provided by the survey. In particular, it considered multi-cohort survey on graduates of the years 1982, 1992, and 2002 for the empirical analysis. As a result of empirical analysis, the wage premium of a major-job match was identified. This result was consistent after the consideration of a sample selection bias and also after modeling the major-job match variable as polychotomously selective. Through an analysis classified by the major, this study identified a relatively high wage premium among Social Science, Engineering, and Science majors. However, there was a difference in the effect of selection among these majors. Also, by assessing cohort effects this study found that the skill mismatch had rapidly progressed in 1992, while difference between 1992 and 2002 cohorts are insignificant. The analysis suggests that wage level is better understood within the context of both sample selection and major-job match, and regardless of model specification the major-job match affects wage strongly.

  • PDF

Logistic Regressions with Sensory Evaluation Data about Hanwoo Steer Beef (한우 거세우 고기 관능평가 데이터의 로지스틱 회귀분석)

  • Lee, Hye-Jung;Kim, Jae-Hee
    • The Korean Journal of Applied Statistics
    • /
    • v.23 no.5
    • /
    • pp.857-870
    • /
    • 2010
  • This study was conducted to investigate the relationship between the socio-demographic factors and the Korean consumers palatability evaluation grades with Hanwoo sensory evaluation data from 2006 to 2008 by National Institute of Animal Science. The dichotomy logistic regression model and the multinomial logistic regression model are fitted with the independent variables such as the consumer living location, age, gender occupation, monthly income, beef cut and the the palatability grade as the categorical dependent variable and tenderness, 리avor and juiciness as the continuous dependent variable. Stepwise variable selection procedure is incorporated to find the final model and odds ratios are calculated to nd the associations between categories.

A Systematic Review on Web 2.0 Adoption (웹2.0 활용 및 도입에 관한 체계적 문헌연구)

  • Kim, Tack-Hyun;Lim, Joa-Sang;Jung, Chul-Yong
    • 한국IT서비스학회:학술대회논문집
    • /
    • 2008.11a
    • /
    • pp.345-348
    • /
    • 2008
  • 본 연구는 체계적인 문헌연구를 통해 웹2.0의 사용이익과 도입요인을 분석하였다. 전자저널 검색결과 259편의 문헌 중에서 도입 기업사례 및 사용요인 관련 연구 10편을 선별하였다. 선택된 논문을 웹2.0기술의 사용이익과 성과, 웹2.0기술 수용요인, 블로그 사용자 행동 및 동기의 주제로 나누어 구분하였다. 이를 통해 본 논문에서는 웹2.0기술에 대한 실제 사용이익을 정보이익, 사회적 이익, 업무관련이익, 지식관련이익으로 제안하였다. 수용요인에 대한 연구결과로는 웹2.0 특성변수를 반영한 요인으로 지각된 참여성과 동시성, 플로우 경험, 지식의 자기효능감, 개인성과기대가 사용의도에 영향을 미치는 변수임을 밝혀내었다.

  • PDF

A study on equating method based on regression analysis (회귀분석에 기초한 균등화 방법에 관한 연구)

  • Cho, Jang-Sik
    • Journal of the Korean Data and Information Science Society
    • /
    • v.21 no.3
    • /
    • pp.513-521
    • /
    • 2010
  • Most of universities have carried out course evaluation to apply the performance appraisal for professor. But, course evaluation depends on characteristics of each class such as class size, type of lecture, evaluator's grade and so on. As the results, such characteristics of each class lead to serious bias which makes lecturers distrust the course evaluation results. Hence, we propose a equating method for the course evaluation by regression analysis which use stepwise variable selection. And we compare proposed method with the other method by Cho et al. (2009) with respect to efficiencies. Also we give the example to which the method is applied.