• Title/Summary/Keyword: Machine Learning & Training

Search Result 819, Processing Time 0.022 seconds

A Study On User Skin Color-Based Foundation Color Recommendation Method Using Deep Learning (딥러닝을 이용한 사용자 피부색 기반 파운데이션 색상 추천 기법 연구)

  • Jeong, Minuk;Kim, Hyeonji;Gwak, Chaewon;Oh, Yoosoo
    • Journal of Korea Multimedia Society
    • /
    • v.25 no.9
    • /
    • pp.1367-1374
    • /
    • 2022
  • In this paper, we propose an automatic cosmetic foundation recommendation system that suggests a good foundation product based on the user's skin color. The proposed system receives and preprocesses user images and detects skin color with OpenCV and machine learning algorithms. The system then compares the performance of the training model using XGBoost, Gradient Boost, Random Forest, and Adaptive Boost (AdaBoost), based on 550 datasets collected as essential bestsellers in the United States. Based on the comparison results, this paper implements a recommendation system using the highest performing machine learning model. As a result of the experiment, our system can effectively recommend a suitable skin color foundation. Thus, our system model is 98% accurate. Furthermore, our system can reduce the selection trials of foundations against the user's skin color. It can also save time in selecting foundations.

Analysis on the Accuracy of Building Construction Cost Estimation by Activation Function and Training Model Configuration (활성화함수와 학습노드 진행 변화에 따른 건축 공사비 예측성능 분석)

  • Lee, Ha-Neul;Yun, Seok-Heon
    • Journal of KIBIM
    • /
    • v.12 no.2
    • /
    • pp.40-48
    • /
    • 2022
  • It is very important to accurately predict construction costs in the early stages of the construction project. However, it is difficult to accurately predict construction costs with limited information from the initial stage. In recent years, with the development of machine learning technology, it has become possible to predict construction costs more accurately than before only with schematic construction characteristics. Based on machine learning technology, this study aims to analyze plans to more accurately predict construction costs by using only the factors influencing construction costs. To the end of this study, the effect of the error rate according to the activation function and the node configuration of the hidden layer was analyzed.

A DDoS attack Mitigation in IoT Communications Using Machine Learning

  • Hailye Tekleselase
    • International Journal of Computer Science & Network Security
    • /
    • v.24 no.4
    • /
    • pp.170-178
    • /
    • 2024
  • Through the growth of the fifth-generation networks and artificial intelligence technologies, new threats and challenges have appeared to wireless communication system, especially in cybersecurity. And IoT networks are gradually attractive stages for introduction of DDoS attacks due to integral frailer security and resource-constrained nature of IoT devices. This paper emphases on detecting DDoS attack in wireless networks by categorizing inward network packets on the transport layer as either "abnormal" or "normal" using the integration of machine learning algorithms knowledge-based system. In this paper, deep learning algorithms and CNN were autonomously trained for mitigating DDoS attacks. This paper lays importance on misuse based DDOS attacks which comprise TCP SYN-Flood and ICMP flood. The researcher uses CICIDS2017 and NSL-KDD dataset in training and testing the algorithms (model) while the experimentation phase. accuracy score is used to measure the classification performance of the four algorithms. the results display that the 99.93 performance is recorded.

A Study on automatic assignment of descriptors using machine learning (기계학습을 통한 디스크립터 자동부여에 관한 연구)

  • Kim, Pan-Jun
    • Journal of the Korean Society for information Management
    • /
    • v.23 no.1 s.59
    • /
    • pp.279-299
    • /
    • 2006
  • This study utilizes various approaches of machine learning in the process of automatically assigning descriptors to journal articles. The effectiveness of feature selection and the size of training set were examined, after selecting core journals in the field of information science and organizing test collection from the articles of the past 11 years. Regarding feature selection, after reducing the feature set using $x^2$ statistics(CHI) and criteria that prefer high-frequency features(COS, GSS, JAC), the trained Support Vector Machines(SVM) performed the best. With respect to the size of the training set, it significantly influenced the performance of Support Vector Machines(SVM) and Voted Perceptron(VTP). However, it had little effect on Naive Bayes(NB).

Machine Learning Based Automatic Categorization Model for Text Lines in Invoice Documents

  • Shin, Hyun-Kyung
    • Journal of Korea Multimedia Society
    • /
    • v.13 no.12
    • /
    • pp.1786-1797
    • /
    • 2010
  • Automatic understanding of contents in document image is a very hard problem due to involvement with mathematically challenging problems originated mainly from the over-determined system induced by document segmentation process. In both academic and industrial areas, there have been incessant and various efforts to improve core parts of content retrieval technologies by the means of separating out segmentation related issues using semi-structured document, e.g., invoice,. In this paper we proposed classification models for text lines on invoice document in which text lines were clustered into the five categories in accordance with their contents: purchase order header, invoice header, summary header, surcharge header, purchase items. Our investigation was concentrated on the performance of machine learning based models in aspect of linear-discriminant-analysis (LDA) and non-LDA (logic based). In the group of LDA, na$\"{\i}$ve baysian, k-nearest neighbor, and SVM were used, in the group of non LDA, decision tree, random forest, and boost were used. We described the details of feature vector construction and the selection processes of the model and the parameter including training and validation. We also presented the experimental results of comparison on training/classification error levels for the models employed.

Machine Learning Based State of Health Prediction Algorithm for Batteries Using Entropy Index (엔트로피 지수를 이용한 기계학습 기반의 배터리의 건강 상태 예측 알고리즘)

  • Sangjin, Kim;Hyun-Keun, Lim;Byunghoon, Chang;Sung-Min, Woo
    • Journal of IKEEE
    • /
    • v.26 no.4
    • /
    • pp.531-536
    • /
    • 2022
  • In order to efficeintly manage a battery, it is important to accurately estimate and manage the SOH(State of Health) and RUL(Remaining Useful Life) of the batteries. Even if the batteries are of the same type, the characteristics such as facility capacity and voltage are different, and when the battery for the training model and the battery for prediction through the model are different, there is a limit to measuring the accuracy. In this paper, We proposed the entropy index using voltage distribution and discharge time is generalized, and four batteries are defined as a training set and a test set alternately one by one to predict the health status of batteries through linear regression analysis of machine learning. The proposed method showed a high accuracy of more than 95% using the MAPE(Mean Absolute Percentage Error).

Development of a Model to Predict the Number of Visitors to Local Festivals Using Machine Learning (머신러닝을 활용한 지역축제 방문객 수 예측모형 개발)

  • Lee, In-Ji;Yoon, Hyun Shik
    • The Journal of Information Systems
    • /
    • v.29 no.3
    • /
    • pp.35-52
    • /
    • 2020
  • Purpose Local governments in each region actively hold local festivals for the purpose of promoting the region and revitalizing the local economy. Existing studies related to local festivals have been actively conducted in tourism and related academic fields. Empirical studies to understand the effects of latent variables on local festivals and studies to analyze the regional economic impacts of festivals occupy a large proportion. Despite of practical need, since few researches have been conducted to predict the number of visitors, one of the criteria for evaluating the performance of local festivals, this study developed a model for predicting the number of visitors through various observed variables using a machine learning algorithm and derived its implications. Design/methodology/approach For a total of 593 festivals held in 2018, 6 variables related to the region considering population size, administrative division, and accessibility, and 15 variables related to the festival such as the degree of publicity and word of mouth, invitation singer, weather and budget were set for the training data in machine learning algorithm. Since the number of visitors is a continuous numerical data, random forest, Adaboost, and linear regression that can perform regression analysis among the machine learning algorithms were used. Findings This study confirmed that a prediction of the number of visitors to local festivals is possible using a machine learning algorithm, and the possibility of using machine learning in research in the tourism and related academic fields, including the study of local festivals, was captured. From a practical point of view, the model developed in this study is used to predict the number of visitors to the festival to be held in the future, so that the festival can be evaluated in advance and the demand for related facilities, etc. can be utilized. In addition, the RReliefF rank result can be used. Considering this, it will be possible to improve the existing local festivals or refer to the planning of a new festival.

IPMN-LEARN: A linear support vector machine learning model for predicting low-grade intraductal papillary mucinous neoplasms

  • Yasmin Genevieve Hernandez-Barco;Dania Daye;Carlos F. Fernandez-del Castillo;Regina F. Parker;Brenna W. Casey;Andrew L. Warshaw;Cristina R. Ferrone;Keith D. Lillemoe;Motaz Qadan
    • Annals of Hepato-Biliary-Pancreatic Surgery
    • /
    • v.27 no.2
    • /
    • pp.195-200
    • /
    • 2023
  • Backgrounds/Aims: We aimed to build a machine learning tool to help predict low-grade intraductal papillary mucinous neoplasms (IPMNs) in order to avoid unnecessary surgical resection. IPMNs are precursors to pancreatic cancer. Surgical resection remains the only recognized treatment for IPMNs yet carries some risks of morbidity and potential mortality. Existing clinical guidelines are imperfect in distinguishing low-risk cysts from high-risk cysts that warrant resection. Methods: We built a linear support vector machine (SVM) learning model using a prospectively maintained surgical database of patients with resected IPMNs. Input variables included 18 demographic, clinical, and imaging characteristics. The outcome variable was the presence of low-grade or high-grade IPMN based on post-operative pathology results. Data were divided into a training/validation set and a testing set at a ratio of 4:1. Receiver operating characteristics analysis was used to assess classification performance. Results: A total of 575 patients with resected IPMNs were identified. Of them, 53.4% had low-grade disease on final pathology. After classifier training and testing, a linear SVM-based model (IPMN-LEARN) was applied on the validation set. It achieved an accuracy of 77.4%, with a positive predictive value of 83%, a specificity of 72%, and a sensitivity of 83% in predicting low-grade disease in patients with IPMN. The model predicted low-grade lesions with an area under the curve of 0.82. Conclusions: A linear SVM learning model can identify low-grade IPMNs with good sensitivity and specificity. It may be used as a complement to existing guidelines to identify patients who could avoid unnecessary surgical resection.

Semi-supervised regression based on support vector machine

  • Seok, Kyungha
    • Journal of the Korean Data and Information Science Society
    • /
    • v.25 no.2
    • /
    • pp.447-454
    • /
    • 2014
  • In many practical machine learning and data mining applications, unlabeled training examples are readily available but labeled ones are fairly expensive to obtain. Therefore semi-supervised learning algorithms have attracted much attentions. However, previous research mainly focuses on classication problems. In this paper, a semi-supervised regression method based on support vector regression (SVR) formulation that is proposed. The estimator is easily obtained via the dual formulation of the optimization problem. The experimental results with simulated and real data suggest superior performance of the our proposed method compared with standard SVR.

A Sparse Data Preprocessing Using Support Vector Regression (Support Vector Regression을 이용한 희소 데이터의 전처리)

  • Jun, Sung-Hae;Park, Jung-Eun;Oh, Kyung-Whan
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.14 no.6
    • /
    • pp.789-792
    • /
    • 2004
  • In various fields as web mining, bioinformatics, statistical data analysis, and so forth, very diversely missing values are found. These values make training data to be sparse. Largely, the missing values are replaced by predicted values using mean and mode. We can used the advanced missing value imputation methods as conditional mean, tree method, and Markov Chain Monte Carlo algorithm. But general imputation models have the property that their predictive accuracy is decreased according to increase the ratio of missing in training data. Moreover the number of available imputations is limited by increasing missing ratio. To settle this problem, we proposed statistical learning theory to preprocess for missing values. Our statistical learning theory is the support vector regression by Vapnik. The proposed method can be applied to sparsely training data. We verified the performance of our model using the data sets from UCI machine learning repository.