Search | Korea Science

Investigating the performance of different decomposition methods in rainfall prediction from LightGBM algorithm

Narimani, Roya;Jun, Changhyun;Nezhad, Somayeh Moghimi;Parisouj, Peiman
- Proceedings of the Korea Water Resources Association Conference
- /
- 2022.05a
- /
- pp.150-150
- /
- 2022
This study investigates the roles of decomposition methods on high accuracy in daily rainfall prediction from light gradient boosting machine (LightGBM) algorithm. Here, empirical mode decomposition (EMD) and singular spectrum analysis (SSA) methods were considered to decompose and reconstruct input time series into trend terms, fluctuating terms, and noise components. The decomposed time series from EMD and SSA methods were used as input data for LightGBM algorithm in two hybrid models, including empirical mode-based light gradient boosting machine (EMDGBM) and singular spectrum analysis-based light gradient boosting machine (SSAGBM), respectively. A total of four parameters (i.e., temperature, humidity, wind speed, and rainfall) at a daily scale from 2003 to 2017 is used as input data for daily rainfall prediction. As results from statistical performance indicators, it indicates that the SSAGBM model shows a better performance than the EMDGBM model and the original LightGBM algorithm with no decomposition methods. It represents that the accuracy of LightGBM algorithm in rainfall prediction was improved with the SSA method when using multivariate dataset.
PDF

Predicting of the Severity of Car Traffic Accidents on a Highway Using Light Gradient Boosting Model (LightGBM 알고리즘을 활용한 고속도로 교통사고심각도 예측모델 구축)

Lee, Hyun-Mi;Jeon, Gyo-Seok;Jang, Jeong-Ah
- The Journal of the Korea institute of electronic communication sciences
- /
- v.15 no.6
- /
- pp.1123-1130
- /
- 2020
This study aims to classify the severity in car crashes using five classification learning models. The dataset used in this study contains 21,013 vehicle crashes, obtained from Korea Expressway Corporation, between the year of 2015-2017 and the LightGBM(Light Gradient Boosting Model) performed well with the highest accuracy. LightGBM, the number of involved vehicles, type of accident, incident location, incident lane type, types of accidents, types of vehicles involved in accidents were shown as priority factors. Based on the results of this model, the establishment of a management strategy for response of highway traffic accident should be presented through a consistent prediction process of accident severity level. This study identifies applicability of Machine Learning Models for Predicting of the Severity of Car Traffic Accidents on a Highway and suggests that various machine learning techniques based on big data that can be used in the future.
https://doi.org/10.13067/JKIECS.2020.15.6.1123 인용 PDF KSCI

A LightGBM and XGBoost Learning Method for Postoperative Critical Illness Key Indicators Analysis

Lei Han;Yiziting Zhu;Yuwen Chen;Guoqiong Huang;Bin Yi
- KSII Transactions on Internet and Information Systems (TIIS)
- /
- v.17 no.8
- /
- pp.2016-2029
- /
- 2023
Accurate prediction of critical illness is significant for ensuring the lives and health of patients. The selection of indicators affects the real-time capability and accuracy of the prediction for critical illness. However, the diversity and complexity of these indicators make it difficult to find potential connections between them and critical illnesses. For the first time, this study proposes an indicator analysis model to extract key indicators from the preoperative and intraoperative clinical indicators and laboratory results of critical illnesses. In this study, preoperative and intraoperative data of heart failure and respiratory failure are used to verify the model. The proposed model processes the datum and extracts key indicators through four parts. To test the effectiveness of the proposed model, the key indicators are used to predict the two critical illnesses. The classifiers used in the prediction are light gradient boosting machine (LightGBM) and eXtreme Gradient Boosting (XGBoost). The predictive performance using key indicators is better than that using all indicators. In the prediction of heart failure, LightGBM and XGBoost have sensitivities of 0.889 and 0.892, and specificities of 0.939 and 0.937, respectively. For respiratory failure, LightGBM and XGBoost have sensitivities of 0.709 and 0.689, and specificity of 0.936 and 0.940, respectively. The proposed model can effectively analyze the correlation between indicators and postoperative critical illness. The analytical results make it possible to find the key indicators for postoperative critical illnesses. This model is meaningful to assist doctors in extracting key indicators in time and improving the reliability and efficiency of prediction.
https://doi.org/10.3837/tiis.2023.08.003 인용 PDF HTML

A Comparative Analysis of Ensemble Learning-Based Classification Models for Explainable Term Deposit Subscription Forecasting (설명 가능한 정기예금 가입 여부 예측을 위한 앙상블 학습 기반 분류 모델들의 비교 분석)

Shin, Zian;Moon, Jihoon;Rho, Seungmin
- The Journal of Society for e-Business Studies
- /
- v.26 no.3
- /
- pp.97-117
- /
- 2021
Predicting term deposit subscriptions is one of representative financial marketing in banks, and banks can build a prediction model using various customer information. In order to improve the classification accuracy for term deposit subscriptions, many studies have been conducted based on machine learning techniques. However, even if these models can achieve satisfactory performance, utilizing them is not an easy task in the industry when their decision-making process is not adequately explained. To address this issue, this paper proposes an explainable scheme for term deposit subscription forecasting. For this, we first construct several classification models using decision tree-based ensemble learning methods, which yield excellent performance in tabular data, such as random forest, gradient boosting machine (GBM), extreme gradient boosting (XGB), and light gradient boosting machine (LightGBM). We then analyze their classification performance in depth through 10-fold cross-validation. After that, we provide the rationale for interpreting the influence of customer information and the decision-making process by applying Shapley additive explanation (SHAP), an explainable artificial intelligence technique, to the best classification model. To verify the practicality and validity of our scheme, experiments were conducted with the bank marketing dataset provided by Kaggle; we applied the SHAP to the GBM and LightGBM models, respectively, according to different dataset configurations and then performed their analysis and visualization for explainable term deposit subscriptions.
https://doi.org/10.7838/jsebs.2021.26.3.097 인용 PDF KSCI

Attack Detection and Classification Method Using PCA and LightGBM in MQTT-based IoT Environment (MQTT 기반 IoT 환경에서의 PCA와 LightGBM을 이용한 공격 탐지 및 분류 방안)

Lee Ji Gu;Lee Soo Jin;Kim Young Won
- Convergence Security Journal
- /
- v.22 no.4
- /
- pp.17-24
- /
- 2022
Recently, machine learning-based cyber attack detection and classification research has been actively conducted, achieving a high level of detection accuracy. However, low-spec IoT devices and large-scale network traffic make it difficult to apply machine learning-based detection models in IoT environment. Therefore, In this paper, we propose an efficient IoT attack detection and classification method through PCA(Principal Component Analysis) and LightGBM(Light Gradient Boosting Model) using datasets collected in a MQTT(Message Queuing Telementry Transport) IoT protocol environment that is also used in the defense field. As a result of the experiment, even though the original dataset was reduced to about 15%, the performance was almost similar to that of the original. It also showed the best performance in comparative evaluation with the four dimensional reduction techniques selected in this paper.
https://doi.org/10.33778/kcsa.2022.22.4.017 인용 PDF KSCI

Performance Analysis of Trading Strategy using Gradient Boosting Machine Learning and Genetic Algorithm

Jang, Phil-Sik
- Journal of the Korea Society of Computer and Information
- /
- v.27 no.11
- /
- pp.147-155
- /
- 2022
In this study, we developed a system to dynamically balance a daily stock portfolio and performed trading simulations using gradient boosting and genetic algorithms. We collected various stock market data from stocks listed on the KOSPI and KOSDAQ markets, including investor-specific transaction data. Subsequently, we indexed the data as a preprocessing step, and used feature engineering to modify and generate variables for training. First, we experimentally compared the performance of three popular gradient boosting algorithms in terms of accuracy, precision, recall, and F1-score, including XGBoost, LightGBM, and CatBoost. Based on the results, in a second experiment, we used a LightGBM model trained on the collected data along with genetic algorithms to predict and select stocks with a high daily probability of profit. We also conducted simulations of trading during the period of the testing data to analyze the performance of the proposed approach compared with the KOSPI and KOSDAQ indices in terms of the CAGR (Compound Annual Growth Rate), MDD (Maximum Draw Down), Sharpe ratio, and volatility. The results showed that the proposed strategies outperformed those employed by the Korean stock market in terms of all performance metrics. Moreover, our proposed LightGBM model with a genetic algorithm exhibited competitive performance in predicting stock price movements.
https://doi.org/10.9708/jksci.2022.27.11.147 인용 PDF KSCI HTML

Potential of multispectral imaging for maturity classification and recognition of oriental melon

Seongmin Lee;Kyoung-Chul Kim;Kangjin Lee;Jinhwan Ryu;Youngki Hong;Byeong-Hyo Cho
- Korean Journal of Agricultural Science
- /
- v.50 no.3
- /
- pp.527-538
- /
- 2023
In this study, we aimed to apply multispectral imaging (713 - 920 nm, 10 bands) for maturity classification and recognition of oriental melons grown in hydroponic greenhouses. A total of 20 oriental melons were selected, and time series multispectral imaging of oriental melons was 7 - 9 times for each sample from April 21, 2023, to May 12, 2023. We used several approaches, such as Savitzky-Golay (SG), standard normal variate (SNV), and Combination of SG and SNV (SG + SNV), for pre-processing the multispectral data. As a result, 713 - 759 nm bands were preprocessed with SG for the maturity classification of oriental melons. Additionally, a Light Gradient Boosting Machine (LightGBM) was used to train the recognition model for oriental melon. R² of recognition model were 0.92, 0.91 for the training and validation sets, respectively, and the F-scores were 96.6 and 79.4% for the training and testing sets, respectively. Therefore, multispectral imaging in the range of 713 - 920 nm can be used to classify oriental melons maturity and recognize their fruits.
https://doi.org/10.7744/kjoas.500317 인용 PDF

Performance Comparison of Neural Network and Gradient Boosting Machine for Dropout Prediction of University Students

Hyeon Gyu Kim
- Journal of the Korea Society of Computer and Information
- /
- v.28 no.8
- /
- pp.49-58
- /
- 2023
Dropouts of students not only cause financial loss to the university, but also have negative impacts on individual students and society together. To resolve this issue, various studies have been conducted to predict student dropout using machine learning. This paper presents a model implemented using DNN (Deep Neural Network) and LGBM (Light Gradient Boosting Machine) to predict dropout of university students and compares their performance. The academic record and grade data collected from 20,050 students at A University, a small and medium-sized 4-year university in Seoul, were used for learning. Among the 140 attributes of the collected data, only the attributes with a correlation coefficient of 0.1 or higher with the attribute indicating dropout were extracted and used for learning. As learning algorithms, DNN (Deep Neural Network) and LightGBM (Light Gradient Boosting Machine) were used. Our experimental results showed that the F1-scores of DNN and LGBM were 0.798 and 0.826, respectively, indicating that LGBM provided 2.5% better prediction performance than DNN.
https://doi.org/10.9708/jksci.2023.28.08.049 인용 PDF HTML

LightGBM Based Prediction of East Sea Vertical Temperature Profile Using XBT Data (XBT 데이터를 이용한 LightGBM 기반 동해 수직 수온분포 예측)

Kim, Young-Joo;Lee, Soo-Jin
- Proceedings of the Korean Society of Computer Information Conference
- /
- 2022.07a
- /
- pp.27-28
- /
- 2022
최근 우리나라에서도 인공지능 모델을 이용한 수온 예측 관련 연구가 활발히 진행되고 있으나 한반도 주변 해역의 수온 예측 연구에서는 주로 해수면 온도만을 예측하는데 중점을 두고 있다. 본 논문에서는 XBT(eXpendable Bathy-Thermograph) 데이터와 LightGBM(Light Gradient Boosting Model)을 이용하여 잠수함 작전 및 대잠전(Anti Submarine Warfare)에 있어서 군사적으로 중요한 동해의 수직 수온분포를 예측하였다. 동해 특정해역의 해수면부터 수심 200m까지 측정된 XBT 데이터를 이용하여 모델을 학습시키고 성능 평가지표(MAE, MSE, RMSE)와 수직 수온분포 그래프를 통해 예측 정확도를 평가하였다.
PDF

Store Sales Prediction Using Gradient Boosting Model (그래디언트 부스팅 모델을 활용한 상점 매출 예측)

Choi, Jaeyoung;Yang, Heeyoon;Oh, Hayoung
- Journal of the Korea Institute of Information and Communication Engineering
- /
- v.25 no.2
- /
- pp.171-177
- /
- 2021
Through the rapid developments in machine learning, there have been diverse utilization approaches not only in industrial fields but also in daily life. Implementations of machine learning on financial data, also have been of interest. Herein, we employ machine learning algorithms to store sales data and present future applications for fintech enterprises. We utilize diverse missing data processing methods to handle missing data and apply gradient boosting machine learning algorithms; XGBoost, LightGBM, CatBoost to predict the future revenue of individual stores. As a result, we found that using median imputation onto missing data with the appliance of the xgboost algorithm has the best accuracy. By employing the proposed method, fintech enterprises and customers can attain benefits. Stores can benefit by receiving financial assistance beforehand from fintech companies, while these corporations can benefit by offering financial support to these stores with low risk.
https://doi.org/10.6109/jkiice.2021.25.2.171 인용 PDF KSCI

Search Result 23, Processing Time 0.026 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)