• Title/Summary/Keyword: principal component regression

Search Result 251, Processing Time 0.032 seconds

Water Quality Assessment and Turbidity Prediction Using Multivariate Statistical Techniques: A Case Study of the Cheurfa Dam in Northwestern Algeria

  • ADDOUCHE, Amina;RIGHI, Ali;HAMRI, Mehdi Mohamed;BENGHAREZ, Zohra;ZIZI, Zahia
    • Applied Chemistry for Engineering
    • /
    • v.33 no.6
    • /
    • pp.563-573
    • /
    • 2022
  • This work aimed to develop a new equation for turbidity (Turb) simulation and prediction using statistical methods based on principal component analysis (PCA) and multiple linear regression (MLR). For this purpose, water samples were collected monthly over a five year period from Cheurfa dam, an important reservoir in Northwestern Algeria, and analyzed for 12 parameters, including temperature (T°), pH, electrical conductivity (EC), turbidity (Turb), dissolved oxygen (DO), ammonium (NH4+), nitrate (NO3-), nitrite (NO2-), phosphate (PO43-), total suspended solids (TSS), biochemical oxygen demand (BOD5) and chemical oxygen demand (COD). The results revealed a strong mineralization of the water and low dissolved oxygen (DO) content during the summer period. High levels of TSS and Turb were recorded during rainy periods. In addition, water was charged with phosphate (PO43-) in the whole period of study. The PCA results revealed ten factors, three of which were significant (eigenvalues >1) and explained 75.5% of the total variance. The F1 and F2 factors explained 36.5% and 26.7% of the total variance, respectively and indicated anthropogenic pollution of domestic agricultural and industrial origin. The MLR turbidity simulation model exhibited a high coefficient of determination (R2 = 92.20%), indicating that 92.20% of the data variability can be explained by the model. TSS, DO, EC, NO3-, NO2-, and COD were the most significant contributing parameters (p values << 0.05) in turbidity prediction. The present study can help with decision-making on the management and monitoring of the water quality of the dam, which is the primary source of drinking water in this region.

Use of the Quantitatively Transformed Field Soil Structure Description of the US National Pedon Characterization Database to Improve Soil Pedotransfer Function

  • Yoon, Sung-Won;Gimenez, Daniel;Nemes, Attila;Chun, Hyen-Chung;Zhang, Yong-Seon;Sonn, Yeon-Kyu;Kang, Seong-Soo;Kim, Myung-Sook;Kim, Yoo-Hak;Ha, Sang-Keun
    • Korean Journal of Soil Science and Fertilizer
    • /
    • v.44 no.5
    • /
    • pp.944-958
    • /
    • 2011
  • Soil hydraulic properties such as hydraulic conductivity or water retention which are costly to measure can be indirectly generated by soil pedotransfer function (PTF) using easily obtainable soil data. The field soil structure description which is routinely recorded could also be used in PTF as an input to reduce the uncertainty. The purposes of this study were to use qualitative morphological soil structure descriptions and soil structural index into PTF and to evaluate their contribution in the prediction of soil hydraulic properties. We transformed categorical morphological descriptions of soil structure into quantitative values using categorical principal component analysis (CATPCA). This approach was tested with a large data set from the US National Pedon Characterization database with the aid of a categorical regression tree analysis. Six different PTFs were used to predict the saturated hydraulic conductivity and those results were averaged to quantify the uncertainty. Quantified morphological description was successively used in multiple linear regression approach to predict the averaged ensemble saturated conductivity. The selected stepwise regression model with only the transformed morphological variables and structural index as predictors predicted the $K_{sat}$ with $r^2$ = 0.48 (p = 0.018), indicating the feasibility of CATPCA approach. In a regression tree analysis, soil structure index and soil texture turned out to be important factors in the prediction of the hydraulic properties. Among structural descriptions size class turned out to be an important grouping parameter in the regression tree. Bulk density, clay content, W33 and structural index explained clusters selected by a two step clustering technique, implying the morphologically described soil structural features are closely related to soil physical as well as hydraulic properties. Although this study provided relatively new method which related soil structure description to soil structure index, the same approach should be tested using a datasets containing the actual measurement of hydraulic properties. More insight on the predictive power of soil structure index to estimate hydraulic properties would be achieved by considering measured the saturated hydraulic conductivity and the soil water retention.

Chemical Oxygen Demand (COD) Model for the Assessment of Water Quality in the Han River, Korea (한강수질 평가를 위한 COD (화학적 산소 요구량) 모델 평가)

  • Kim, Jae Hyoun;Jo, Jinnam
    • Journal of Environmental Health Sciences
    • /
    • v.42 no.4
    • /
    • pp.280-292
    • /
    • 2016
  • Objectives: The objective of this study was to build COD regression models for the Han River and evaluate water quality. Methods: Water quality data sets for the dry season (as of January) during a four-year period (2012-2015) were collected from the database of the Han River automatic water quality monitoring stations. Statistical techniques, including combined genetic algorithm-multiple linear regression (GA-MLR) were used to build five-descriptor COD models. Multivariate statistical techniques such as principal component analysis (PCA) and cluster analysis (CA) are useful tools for extracting meaningful information. Results: The $r^2$ of the best COD models provided significant high values (> 0.8) between 2012 and 2015. Total organic carbon (TOC) was a surrogate indicator for COD (as COD/TOC) with high reliability ($r^2=0.63$ in 2012, $r^2=0.75$ for 2013, $r^2=0.79$ for 2014 and $r^2=0.85$ for 2015). The ratios of COD/TOC were calculated as 2.08 in 2012, 1.79 in 2013, 1.52 and 1.45 in 2015, indicating that biodegradability in the water body of the Han River was being sustained, thereby further improving water quality. The BOD/COD ratio supported these findings. The cluster analysis revealed higher annual levels of microorganisms and phosphorous at stations along the Hangang-Seoul and Hantangang areas. Nevertheless, the overall water quality over the last four years showed an observable trend toward continuous improvement. These findings also suggest that non-point pollution control strategies should consider the influence of upstreams and downstreams to protect water quality in the Han River. Conclusion: This data analysis procedure provided an efficient and comprehensive tool to interpret complex water quality data matrices. Results from a trend analysis provided much important information about sources and parameters for Han River water quality management.

Evaluation of a Traditional Korean Medicine Content Factor and Satisfaction with the Drama "Daejanggeum" (드라마 "대장금"의 한의학 콘텐츠 요소 및 만족도 평가)

  • Kim, Song-Yi;Kim, Ho-Sun;Nam, Min-Ho;Li, Yuejuan;Chung, Hung-Chiang;Park, Hi-Joon;Lee, Hye-Jung;Chae, Youn-Byoung
    • Journal of Acupuncture Research
    • /
    • v.27 no.1
    • /
    • pp.11-20
    • /
    • 2010
  • Objectives : The study was performed to evaluate a traditional Korean Medicine content in drama "Daejanggeum". Methods : One hundred sixty-nine participants in Taiwan responded to the survey with 10 items, regarding components of success of drama "Daejanggeum". Principal component factor analysis and multiple regression analysis were performed to identify the possible factors to satisfaction with watching drama "Daejanggeum". Results : Factor analysis revealed that dramatic factor(44.8%), content factor(12.3%), and cultural factor(11.3%) were the most important factors to success of drama "Daejanggeum". Multiple regression analysis showed that dramatic factor(beta = .342), content factor(beta = .278), and cultural factor(beta = .131) were associated with the satisfaction with watching drama "Daejanggeum"($R^2$ = .394, with F = 32.280, p<.001). Conclusions : This study demonstrated that dramatic factor, content factor, and cultural factor are the most important factors associated with satisfaction with drama "Daejanggeum" in Taiwan. These findings suggest that a traditional Korean Medicine as a content factor would be very influential in enhancing the possibility of success of drama.

An Analysis of the Economic Effects of R&D Investment in the IT Industry (IT산업 연구개발 투자의 경제적 효과 분석)

  • Hong, Jae-Pyo;Choi, Na-Lin;Kim, Pang-Ryong
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.37B no.9
    • /
    • pp.837-848
    • /
    • 2012
  • This study has conducted the economic effects of R&D investment in the IT industry using multi-regression analysis with three independent variables; capital stock, labor input and R&D stock. In this study, the IT industry has been categorized into three sub-industries; broadcasting communication appliances, information appliances and electronic components industry. Our analysis has found that auto-correlation shows considerable levels whereas figures of t-value and R-square show significant levels among all the IT sub-industries. Meanwhile, the values of R&D stock in the information appliances industry and that of labor input coefficients in the electronic components industry were minus, thus multi-collinearity was suspected. We have solved the problems regarding auto-correlation and multi-collinearity through Cochrane-Orcutt estimation and principal components analysis. This paper has derived the implications that R&D investment in the broadcasting communication industry is much more influential than any other IT sub-industry.

Variable Selection for Multi-Purpose Multivariate Data Analysis (다목적 다변량 자료분석을 위한 변수선택)

  • Huh, Myung-Hoe;Lim, Yong-Bin;Lee, Yong-Goo
    • The Korean Journal of Applied Statistics
    • /
    • v.21 no.1
    • /
    • pp.141-149
    • /
    • 2008
  • Recently we frequently analyze multivariate data with quite large number of variables. In such data sets, virtually duplicated variables may exist simultaneously even though they are conceptually distinguishable. Duplicate variables may cause problems such as the distortion of principal axes in principal component analysis and factor analysis and the distortion of the distances between observations, i.e. the input for cluster analysis. Also in supervised learning or regression analysis, duplicated explanatory variables often cause the instability of fitted models. Since real data analyses are aimed often at multiple purposes, it is necessary to reduce the number of variables to a parsimonious level. The aim of this paper is to propose a practical algorithm for selection of a subset of variables from a given set of p input variables, by the criterion of minimum trace of partial variances of unselected variables unexplained by selected variables. The usefulness of proposed method is demonstrated in visualizing the relationship between selected and unselected variables, in building a predictive model with very large number of independent variables, and in reducing the number of variables and purging/merging categories in categorical data.

Discrimination of African Yams Containing High Functional Compounds Using FT-IR Fingerprinting Combined by Multivariate Analysis and Quantitative Prediction of Functional Compounds by PLS Regression Modeling (FT-IR 스펙트럼 데이터의 다변량 통계분석을 이용한 고기능성 아프리칸 얌 식별 및 기능성 성분 함량 예측 모델링)

  • Song, Seung Yeob;Jie, Eun Yee;Ahn, Myung Suk;Kim, Dong Jin;Kim, In Jung;Kim, Suk Weon
    • Horticultural Science & Technology
    • /
    • v.32 no.1
    • /
    • pp.105-114
    • /
    • 2014
  • We established a high throughput screening system of African yam tuber lines which contain high contents of total carotenoids, flavonoids, and phenolic compounds using ultraviolet-visible (UV-VIS) spectroscopy and Fourier transform infrared (FT-IR) spectroscopy in combination with multivariate analysis. The total carotenoids contents from 62 African yam tubers varied from 0.01 to $0.91{\mu}g{\cdot}g^{-1}$ dry weight (wt). The total flavonoids and phenolic compounds also varied from 12.9 to $229{\mu}g{\cdot}g^{-1}$ and from 0.29 to $5.2mg{\cdot}g^{-1}$dry wt. FT-IR spectra confirmed typical spectral differences between the frequency regions of 1,700-1,500, 1,500-1,300 and $1,100-950cm^{-1}$, respectively. These spectral regions were reflecting the quantitative and qualitative variations of amide I, II from amino acids and proteins ($1,700-1,500cm^{-1}$), phosphodiester groups from nucleic acid and phospholipid ($1,500-1,300cm^{-1}$) and carbohydrate compounds ($1,100-950cm^{-1}$). Principal component analysis (PCA) and subsequent partial least square-discriminant analysis (PLS-DA) were able to discriminate the 62 African yam tuber lines into three separate clusters corresponding to their taxonomic relationship. The quantitative prediction modeling of total carotenoids, flavonoids, and phenolic compounds from African yam tuber lines were established using partial least square regression algorithm from FT-IR spectra. The regression coefficients ($R^2$) between predicted values and estimated values of total carotenoids, flavonoids and phenolic compounds were 0.83, 0.86, and 0.72, respectively. These results showed that quantitative predictions of total carotenoids, flavonoids, and phenolic compounds were possible from FT-IR spectra of African yam tuber lines with higher accuracy. Therefore we suggested that quantitative prediction system established in this study could be applied as a rapid selection tool for high yielding African yam lines.

Prediction of golf scores on the PGA tour using statistical models (PGA 투어의 골프 스코어 예측 및 분석)

  • Lim, Jungeun;Lim, Youngin;Song, Jongwoo
    • The Korean Journal of Applied Statistics
    • /
    • v.30 no.1
    • /
    • pp.41-55
    • /
    • 2017
  • This study predicts the average scores of top 150 PGA golf players on 132 PGA Tour tournaments (2013-2015) using data mining techniques and statistical analysis. This study also aims to predict the Top 10 and Top 25 best players in 4 different playoffs. Linear and nonlinear regression methods were used to predict average scores. Stepwise regression, all best subset, LASSO, ridge regression and principal component regression were used for the linear regression method. Tree, bagging, gradient boosting, neural network, random forests and KNN were used for nonlinear regression method. We found that the average score increases as fairway firmness or green height or average maximum wind speed increases. We also found that the average score decreases as the number of one-putts or scrambling variable or longest driving distance increases. All 11 different models have low prediction error when predicting the average scores of PGA Tournaments in 2015 which is not included in the training set. However, the performances of Bagging and Random Forest models are the best among all models and these two models have the highest prediction accuracy when predicting the Top 10 and Top 25 best players in 4 different playoffs.

An intelligent sensor system with reconstruction mechanism of faulty signal

  • Jung, Young-Su;Hyun, Woong-Keun;Yoon, In-Mo;Jung, Young-Kee;Kim, C.S.;Kim, Nam-Ho
    • 제어로봇시스템학회:학술대회논문집
    • /
    • 2003.10a
    • /
    • pp.1231-1234
    • /
    • 2003
  • A sensor working in outdoor may generate some faulty signal owing to dust and high temperature. This paper describes an intelligent sensor system and controller which has a reconstruction mechanism for faulty signal. The faulty signals are dievided into two types as linear distortion and non linear distortion, respectively. The linear distorted signal is due to dust, and non linear distorted signal is due to physical breakdown of sensor or high temperature. These distorted signal have been reconstructed by the proposed method based on polynomial regression method and principal component analysis approach.. The proposed method has been applied to sun tracking system working in outdoor. For a robust and precision control of sun tracker, a fuzzy controller was also proposed. The fuzzy controller controls the tracker by using the collected sensor signal. The tolerance of the position control is within 1.5 degree. To show the validity of the developed system, some experiments in the field were illustrated.

  • PDF

Enhancement of the Virtual Metrology Performance for Plasma-assisted Processes by Using Plasma Information (PI) Parameters

  • Park, Seolhye;Lee, Juyoung;Jeong, Sangmin;Jang, Yunchang;Ryu, Sangwon;Roh, Hyun-Joon;Kim, Gon-Ho
    • Proceedings of the Korean Vacuum Society Conference
    • /
    • 2015.08a
    • /
    • pp.132-132
    • /
    • 2015
  • Virtual metrology (VM) model based on plasma information (PI) parameter for C4F8 plasma-assisted oxide etching processes is developed to predict and monitor the process results such as an etching rate with improved performance. To apply fault detection and classification (FDC) or advanced process control (APC) models on to the real mass production lines efficiently, high performance VM model is certainly required and principal component regression (PCR) is preferred technique for VM modeling despite this method requires many number of data set to obtain statistically guaranteed accuracy. In this study, as an effective method to include the 'good information' representing parameter into the VM model, PI parameters are introduced and applied for the etch rate prediction. By the adoption of PI parameters of b-, q-factors and surface passivation parameters as PCs into the PCR based VM model, information about the reactions in the plasma volume, surface, and sheath regions can be efficiently included into the VM model; thus, the performance of VM is secured even for insufficient data set provided cases. For mass production data of 350 wafers, developed PI based VM (PI-VM) model was satisfied required prediction accuracy of industry in C4F8 plasma-assisted oxide etching process.

  • PDF