• 제목/요약/키워드: Statistical Learning Model

검색결과 545건 처리시간 0.027초

스파크에서 스칼라와 R을 이용한 머신러닝의 비교 (Comparison of Scala and R for Machine Learning in Spark)

  • 류우석
    • 한국전자통신학회논문지
    • /
    • 제18권1호
    • /
    • pp.85-90
    • /
    • 2023
  • 보건의료분야 데이터 분석 방법론이 기존의 통계 중심의 연구방법에서 머신러닝을 이용한 예측 연구로 전환되고 있다. 본 연구에서는 다양한 머신러닝 도구들을 살펴보고, 보건의료분야에서 많이 사용하고 있는 통계 도구인 R을 빅데이터 머신러닝에 적용하기 위해 R과 스파크를 연계한 프로그래밍 모델들을 비교한다. 그리고, R을 스파크 환경에서 수행하는 SparkR을 이용한 선형회귀모델 학습의 성능을 스파크의 기본 언어인 스칼라를 이용한 모델과 비교한다. 실험 결과 SparkR을 이용할 때의 학습 수행 시간이 스칼라와 비교하여 10~20% 정도 증가하였다. 결과로 제시된 성능 저하를 감안한다면 기존의 통계분석 도구인 R을 그대로 활용 가능하다는 측면에서 SparkR의 분산 처리의 유용성을 확인하였다.

Semi-supervised learning using similarity and dissimilarity

  • Seok, Kyung-Ha
    • Journal of the Korean Data and Information Science Society
    • /
    • 제22권1호
    • /
    • pp.99-105
    • /
    • 2011
  • We propose a semi-supervised learning algorithm based on a form of regularization that incorporates similarity and dissimilarity penalty terms. Our approach uses a graph-based encoding of similarity and dissimilarity. We also present a model-selection method which employs cross-validation techniques to choose hyperparameters which affect the performance of the proposed method. Simulations using two types of dat sets demonstrate that the proposed method is promising.

면역 알고리즘 기반의 서포트 벡터 회귀를 이용한 소프트웨어 신뢰도 추정 (Estimation of Software Reliability with Immune Algorithm and Support Vector Regression)

  • 권기태;이준길
    • 한국IT서비스학회지
    • /
    • 제8권4호
    • /
    • pp.129-140
    • /
    • 2009
  • The accurate estimation of software reliability is important to a successful development in software engineering. Until recent days, the models using regression analysis based on statistical algorithm and machine learning method have been used. However, this paper estimates the software reliability using support vector regression, a sort of machine learning technique. Also, it finds the best set of optimized parameters applying immune algorithm, changing the number of generations, memory cells, and allele. The proposed IA-SVR model outperforms some recent results reported in the literature.

Predicting Reports of Theft in Businesses via Machine Learning

  • JungIn, Seo;JeongHyeon, Chang
    • International Journal of Advanced Culture Technology
    • /
    • 제10권4호
    • /
    • pp.499-510
    • /
    • 2022
  • This study examines the reporting factors of crime against business in Korea and proposes a corresponding predictive model using machine learning. While many previous studies focused on the individual factors of theft victims, there is a lack of evidence on the reporting factors of crime against a business that serves the public good as opposed to those that protect private property. Therefore, we proposed a crime prevention model for the willingness factor of theft reporting in businesses. This study used data collected through the 2015 Commercial Crime Damage Survey conducted by the Korea Institute for Criminal Policy. It analyzed data from 834 businesses that had experienced theft during a 2016 crime investigation. The data showed a problem with unbalanced classes. To solve this problem, we jointly applied the Synthetic Minority Over Sampling Technique and the Tomek link techniques to the training data. Two prediction models were implemented. One was a statistical model using logistic regression and elastic net. The other involved a support vector machine model, tree-based machine learning models (e.g., random forest, extreme gradient boosting), and a stacking model. As a result, the features of theft price, invasion, and remedy, which are known to have significant effects on reporting theft offences, can be predicted as determinants of such offences in companies. Finally, we verified and compared the proposed predictive models using several popular metrics. Based on our evaluation of the importance of the features used in each model, we suggest a more accurate criterion for predicting var.

Prediction & Assessment of Change Prone Classes Using Statistical & Machine Learning Techniques

  • Malhotra, Ruchika;Jangra, Ravi
    • Journal of Information Processing Systems
    • /
    • 제13권4호
    • /
    • pp.778-804
    • /
    • 2017
  • Software today has become an inseparable part of our life. In order to achieve the ever demanding needs of customers, it has to rapidly evolve and include a number of changes. In this paper, our aim is to study the relationship of object oriented metrics with change proneness attribute of a class. Prediction models based on this study can help us in identifying change prone classes of a software. We can then focus our efforts on these change prone classes during testing to yield a better quality software. Previously, researchers have used statistical methods for predicting change prone classes. But machine learning methods are rarely used for identification of change prone classes. In our study, we evaluate and compare the performances of ten machine learning methods with the statistical method. This evaluation is based on two open source software systems developed in Java language. We also validated the developed prediction models using other software data set in the same domain (3D modelling). The performance of the predicted models was evaluated using receiver operating characteristic analysis. The results indicate that the machine learning methods are at par with the statistical method for prediction of change prone classes. Another analysis showed that the models constructed for a software can also be used to predict change prone nature of classes of another software in the same domain. This study would help developers in performing effective regression testing at low cost and effort. It will also help the developers to design an effective model that results in less change prone classes, hence better maintenance.

연관분석을 이용한 마코프 논리네트워크의 1차 논리 공식 생성과 가중치 학습방법 (First-Order Logic Generation and Weight Learning Method in Markov Logic Network Using Association Analysis)

  • 안길승;허선
    • 산업경영시스템학회지
    • /
    • 제38권1호
    • /
    • pp.74-82
    • /
    • 2015
  • Two key challenges in statistical relational learning are uncertainty and complexity. Standard frameworks for handling uncertainty are probability and first-order logic respectively. A Markov logic network (MLN) is a first-order knowledge base with weights attached to each formula and is suitable for classification of dataset which have variables correlated with each other. But we need domain knowledge to construct first-order logics and a computational complexity problem arises when calculating weights of first-order logics. To overcome these problems we suggest a method to generate first-order logics and learn weights using association analysis in this study.

Genetic classification of various familial relationships using the stacking ensemble machine learning approaches

  • Su Jin Jeong;Hyo-Jung Lee;Soong Deok Lee;Ji Eun Park;Jae Won Lee
    • Communications for Statistical Applications and Methods
    • /
    • 제31권3호
    • /
    • pp.279-289
    • /
    • 2024
  • Familial searching is a useful technique in a forensic investigation. Using genetic information, it is possible to identify individuals, determine familial relationships, and obtain racial/ethnic information. The total number of shared alleles (TNSA) and likelihood ratio (LR) methods have traditionally been used, and novel data-mining classification methods have recently been applied here as well. However, it is difficult to apply these methods to identify familial relationships above the third degree (e.g., uncle-nephew and first cousins). Therefore, we propose to apply a stacking ensemble machine learning algorithm to improve the accuracy of familial relationship identification. Using real data analysis, we obtain superior relationship identification results when applying meta-classifiers with a stacking algorithm rather than applying traditional TNSA or LR methods and data mining techniques.

그래픽 계산기를 활용하는 수학과 교수-학습 자료 모형 개발 연구 (Study on the Development of a Model for Teaching and Learning Mathematics Using Graphic Calculators)

  • 강옥기
    • 대한수학교육학회지:수학교육학연구
    • /
    • 제8권2호
    • /
    • pp.453-474
    • /
    • 1998
  • This study is focused on the possibility if we can use graphic calculators in teaching and learning school mathematics. This study is consisted with four main chapters. In chapter II, the functions of the graphic calculator EL-9600 produced by Sharp Corporation was analyzed focused on the possibilities if the functions could be used in teaching and learning school mathematics. Calculating of real numbers and complex numbers, solving equations and system of linear equations, calculating of matrices, graphing of several functions including polynomial functions, trigonometric functions, exponential and logarithmic functions, calculation of differential and integrals, arranging of statical data, graphing of statistical data, testing of statistical hypotheses, and other more useful functions were founded. In Chapter III, a mathematics textbook developed by Core-Plus Mathematics Project was analyzed focused on how a graphic calculator was used in teaching and learning mathematics, In the textbook, graphic calculator was used as a tool in understanding mathematical concepts and solving problems. Graphic calculator is not just a tool to do complex computations but a tool used in the processes of doing mathematics, In chapter IV, the 7th mathematics curriculum for korean secondary schools was analyzed to find the contents could be taught by using graphic calculators. Most of the domains, except geometric figure, were found that they could be taught by using graphic calculators, In chapter V, a model of a unit using graphic calculator in teaching 7th mathematics curriculum was developed. In this model, graphic calculator was used as a tool in the processes of understanding mathematical concepts and solving problems. This study suggests the possibilities that we can use graphic calculators effectively in teaching and learning mathematical concepts and problem solving for most domains of secondary school mathematics.

  • PDF

대수형 학습효과에 근거한 소프트웨어 신뢰모형에 관한 통계적 공정관리 비교 연구 (The Assessing Comparative Study for Statistical Process Control of Software Reliability Model Based on Logarithmic Learning Effects)

  • 김경수;김희철
    • 디지털융복합연구
    • /
    • 제11권12호
    • /
    • pp.319-326
    • /
    • 2013
  • 소프트웨어의 디버깅 오류의 발생 시간에 의존하는 많은 소프트웨어 신뢰성 모델이 연구되었다. 소프트웨어 오류 탐색 기법은 사전에 알지 못하지만 자동적으로 발견되는 에러를 고려한 영향요인과 사전 경험에 의하여 세밀하게 에러를 발견하기 위하여 테스팅 관리자가 설정해놓은 요인인 학습효과의 특성에 대한 문제를 비교 제시 하였다. 본 연구에서는 학습효과 비동질적인 유한고장모형 분석을 위한 모수 추정은 우도함수를 이용하였다. 소프트웨어 시장에 인도하기 위한 결정에 대하여 조건부 고장률은 중요한 변수가 되고 이러한 고장 모델은 실제 상황에서 많이 사용되고 있다. 통계적 공정 관리 (SPC)는 소프트웨어 오류의 예측을 모니터링 함으로써 소프트웨어의 신뢰성 향상에 크게 기여할 수 있다. 이러한 컨트롤 차트는 널리 소프트웨어 산업의 소프트웨어 프로세스 제어를 위해 사용된다. 본 연구에서는 로그 위험 학습 효과 속성의 비동질적인 포아송 과정의 평균값 기능을 사용한 컨트롤 메커니즘을 제안하였다.

온라인 서포트벡터기계를 이용한 온라인 비정상 사건 탐지 (Online abnormal events detection with online support vector machine)

  • 박혜정
    • Journal of the Korean Data and Information Science Society
    • /
    • 제22권2호
    • /
    • pp.197-206
    • /
    • 2011
  • 신호처리 관련 응용문제에서는 신호에서 실시간으로 발생하는 비정상적인 사건들을 탐지하는 것이 매우 중요하다. 이전에 알려져 있는 비정상 사건 탐지방법들은 신호에 대한 명확한 통계적인 모형을 가정하고, 비정상적인 신호들은 통계적인 모형의 가정 하에서 비정상적인 사건들로 해석한다. 탐지방법으로 최대우도와 베이즈 추정 이론이 많이 사용되고 있다. 그러나 앞에서 언급한 방법으로는 로버스트 하고 다루기 쉬운 모형을 추정한다는 것은 쉽지가 않다. 좀 더 로버스트한 모형을 추정할 수 있는 방법이 필요하다. 본 논문에서는 로버스트 하다고 알려져 있는 서포트 벡터 기계를 이용하여 온라인으로 비정상적인 신호를 탐지하는 방법을 제안한다.