• Title/Summary/Keyword: PLS Regression

검색결과 175건 처리시간 0.024초

Shrinkage Structure of Ridge Partial Least Squares Regression

  • Kim, Jong-Duk
    • Journal of the Korean Data and Information Science Society
    • /
    • 제18권2호
    • /
    • pp.327-344
    • /
    • 2007
  • 다중공선성의 데이터에 사용되는 대표적인 편향회귀방법은 능형회귀(RR), 주성분회귀(PCR), 부분최소제곱회귀(PLS) 등이다. 이 회귀방법들은 계수베거 추정량의 놈(norm)이 모두 보통 최소제곱회귀(OLS)의 추정량의 놈보다 작아진다는 의미에서 축소회귀라 부른다. 새로운 회귀방법으로 RR과 PCR을 결합한 능형주성분회귀(RPCR)가 있고 RR과 PLS를 결합한 능형부분최소제곱회귀(RPLS)가 있으며 이들도 또한 축소회귀이다. 이들 추정량은 X'X의 고유벡터들의 선형결합으로 나타낼 수 있고 따라서 각 고유방향에서 OLS에 비해 얼마나 축소되는지를 연구할 수 있다. 본 논문에서는 먼저 이들 추정량을 일반적인 축소인자의 식으로 나타내고 이를 이용하여 MSE의 일반식을 구하였으며 PLS 추정량의 MSE 식도 구하였다. 그리고 RPLS의 축소인자 식을 두 가지 다른 형태로 유도하였다. RPLS의 경우도 이 축소인자 식을 MSE의 일반식에 대입하면 MSE 식이 바로 얻어진다. 그러나 PLS나 RPLS의 축소인자는 y의 복잡한 비선형이 되어 결정적이 아니므로 이들 추정량의 MSE는 근사적인 식이라 할 수 있다. 따라서 PLS나 RPLS를 평가하기 위해 이 MSE를 사용하는 것은 제한적이며, 경험적인 방법으로 이들 회귀의 수행성을 평가하는 것이 필요하다. 다중공선성의 대표적인 데이터인 근적외선 분광 데이터를 이용하여 이 유도된 회귀의 축소인자 값이 인자수에 따라 어떻게 변화하는지와 전체적인 축소 비율도 살펴보았다. 이들의 축소 형태를 잘 이해하면 회귀방법들의 예측력과 안정성을 파악하는데 많은 도움이 되리라 판단된다.

  • PDF

FT-IR 스펙트럼 데이터의 다변량 통계분석을 이용한 곶감의 원산지 및 품종 식별 (Discrimination of Cultivars and Cultivation Origins from the Sepals of Dry Persimmon Using FT-IR Spectroscopy Combined with Multivariate Analysis)

  • 허설혜;김석원;민병환
    • 한국식품과학회지
    • /
    • 제47권1호
    • /
    • pp.20-26
    • /
    • 2015
  • 본 연구에서는 상업용 곶감의 꽃받침과 종자를 이용하여 대사체 수준에서의 원산지와 품종 식별 체계를 확립하였다. 실험에 이용된 곶감 시료는 국내산 곶감 함안수시(Hamansusi), 예천고종시(Yecheongojongsi), 산청단성시(Sancheongdanseongsi), 그리고 논산월하시(Nonsanwalhasi) 4개 품종과 국내에서 판매되고 있는 중국산 곶감 2개 종류의 꽃받침과 종자를 사용하였으며, 꽃받침과 종자 시료의 전세포 추출물로부터 FT-IR 스펙트럼 데이터를 기반으로 다변량 통계분석(PCA, PLS-DA)을 실시하였다. 이 결과 국내산 곶감 4품종과 중국산 곶감 2종류가 두 그룹으로 확연히 나뉘어지는 것을 확인할 수 있었다. 상업용 곶감의 꽃받침을 PLS regression을 실시한 결과 국내산과 중국산 곶감을 100% 예측할 수 있었다. 또한 곶감 종자를 이용하여 품종 식별한 결과 각 4개의 그룹으로 나뉘어지는 것을 확인할 수 있었으며, PLS regression을 실시한 결과 약 86%의 정확도로 품종 식별이 가능함을 알 수 있었다. FT-IR 스펙트럼 분석의 간편성과 신속성을 고려할 때, 본 연구 결과는 상업용 곶감에 대한 원산지나 품종 식별의 신속한 수단으로 활용할 수 있을 것으로 예상된다. 더 나아가 본 기술을 이용하여 다른 농산물의 원산지 또는 품종 식별 수단으로 활용이 가능할 것으로 기대된다.

정조 상태에서 백미에 대한 완전미율의 비파괴 예측 (Non-Destructive Prediction of Head Rice Ratios using NIR Spectra of Hulled Rice)

  • 권영립;조승현;이재흥;서경원;최동칠
    • 한국작물학회지
    • /
    • 제53권3호
    • /
    • pp.244-250
    • /
    • 2008
  • 도정하지 않은 정조의 81 시료로부터 스펙트럼을 수집하고, 백미 완전미도정수율 예측 희귀모델을 개발하기 위해 검량식을 작성한 결과 스펙트럼을 8 nm 간격으로 지정하고, 1차미분 방법으로 검량식을 작성한 완전미율의 결정계수는 MPLS에서 0.8353, PLS 방법에서 0.8416, PCR에서 0.5277를 나타냈다. 스펙트럼을 20 nm 간격으로 지정하고 1차미분 방법으로 검량식을 작성하였다. 완전미율의 결정계수는 MPLS에서 0.8144, PLS 방법에서 0.8354, PCR에서 0.6809를 나타냈다. 스펙트럼을 8 nm 간격으로 지정하고 2차미분 방법으로 검량식을 작성하였다.완 전미율의 결정계수는 MPLS 방법에서 0.7994, PLS에서 0.8017, PCR에서 0.4473을 나타냈다. 스펙트럼을 20 nm 간격으로 지정하고 2차미분 방법으로 검량식을 작성하였다. 완전미율의 결정계수는MPLS 방법에서 0.8004, PLS에서 0.8493, PCR에서 0.6609을 나타냈다.

MEAT SPECIATION USING A HIERARCHICAL APPROACH AND LOGISTIC REGRESSION

  • Arnalds, Thosteinn;Fearn, Tom;Downey, Gerard
    • 한국근적외분광분석학회:학술대회논문집
    • /
    • 한국근적외분광분석학회 2001년도 NIR-2001
    • /
    • pp.1245-1245
    • /
    • 2001
  • Food adulteration is a serious consumer fraud and a matter of concern to food processors and regulatory agencies. A range of analytical methods have been investigated to facilitate the detection of adulterated or mis-labelled foods & food ingredients but most of these require sophisticated equipment, highly-qualified staff and are time-consuming. Regulatory authorities and the food industry require a screening technique which will facilitate fast and relatively inexpensive monitoring of food products with a high level of accuracy. Near infrared spectroscopy has been investigated for its potential in a number of authenticity issues including meat speciation (McElhinney, Downey & Fearn (1999) JNIRS, 7(3), 145-154; Downey, McElhinney & Fearn (2000). Appl. Spectrosc. 54(6), 894-899). This report describes further analysis of these spectral sets using a hierarchical approach and binary decisions solved using logistic regression. The sample set comprised 230 homogenized meat samples i. e. chicken (55), turkey (54), pork (55), beef (32) and lamb (34) purchased locally as whole cuts of meat over a 10-12 week period. NIR reflectance spectra were recorded over the wavelength range 400-2498nm at 2nm intervals on a NIR Systems 6500 scanning monochromator. The problem was defined as a series of binary decisions i. e. is the meat red or white\ulcorner is the red meat beef or lamb\ulcorner, is the white meat pork or poultry\ulcorner etc. Each of these decisions was made using an individual binary logistic model based on scores derived from principal component or partial least squares (PLS1 and PLS2) analysis. The results obtained were equal to or better than previous reports using factorial discriminant analysis, K-nearest neighbours and PLS2 regression. This new approach using a combination of exploratory and logistic analyses also appears to have advantages of transparency and the use of inherent structure in the spectral data. Additionally, it allows for the use of different data transforms and multivariate regression techniques at each decision step.

  • PDF

MEAT SPECIATION USING A HIERARCHICAL APPROACH AND LOGISTIC REGRESSION

  • Arnalds, Thosteinn;Fearn, Tom;Downey, Gerard
    • 한국근적외분광분석학회:학술대회논문집
    • /
    • 한국근적외분광분석학회 2001년도 NIR-2001
    • /
    • pp.1152-1152
    • /
    • 2001
  • Food adulteration is a serious consumer fraud and a matter of concern to food processors and regulatory agencies. A range of analytical methods have been investigated to facilitate the detection of adulterated or mis-labelled foods & food ingredients but most of these require sophisticated equipment, highly-qualified staff and are time-consuming. Regulatory authorities and the food industry require a screening technique which will facilitate fast and relatively inexpensive monitoring of food products with a high level of accuracy. Near infrared spectroscopy has been investigated for its potential in a number of authenticity issues including meat speciation (McElhinney, Downey & Fearn (1999) JNIRS, 7(3), 145 154; Downey, McElhinney & Fearn (2000). Appl. Spectrosc. 54(6), 894-899). This report describes further analysis of these spectral sets using a hierarchical approach and binary decisions solved using logistic regression. The sample set comprised 230 homogenized meat samples i. e. chicken (55), turkey (54), pork (55), beef (32) and lamb (34) purchased locally as whole cuts of meat over a 10-12 week period. NIR reflectance spectra were recorded over the wavelength range 400-2498nm at 2nm intervals on a NIR Systems 6500 scanning monochromator. The problem was defined as a series of binary decisions i. e. is the meat red or white\ulcorner is the red meat beef or lamb\ulcorner, is the white meat pork or poultry\ulcorner etc. Each of these decisions was made using an individual binary logistic model based on scores derived from principal component or partial least squares (PLS1 and PLS2) analysis. The results obtained were equal to or better than previous reports using factorial discriminant analysis, K-nearest neighbours and PLS2 regression. This new approach using a combination of exploratory and logistic analyses also appears to have advantages of transparency and the use of inherent structure in the spectral data. Additionally, it allows for the use of different data transforms and multivariate regression techniques at each decision step.

  • PDF

PREPROCESSING EFFECTS ON ON-LINE SSC MEASUREMENT OF FUJI APPLE BY NIR SPECTROSCOPY

  • Ryu, D.S.;Noh, S.H.;Hwang, I.G.
    • 한국농업기계학회:학술대회논문집
    • /
    • 한국농업기계학회 2000년도 THE THIRD INTERNATIONAL CONFERENCE ON AGRICULTURAL MACHINERY ENGINEERING. V.III
    • /
    • pp.560-568
    • /
    • 2000
  • The aims of this research were to investigate the preprocessing effect of spectrum data on prediction performance and to develop a robust model to predict SSC in intact apple. Spectrum data of 320 Fuji apples were measured with the on-line transmittance measurement system at the wavelength range of 550∼1100nm. Preprocess methods adopted for the tests were Savitzky Golay, MSC, SNV, first derivative and OSC. Several combinations of those methods were applied to the raw spectrum data set to investigate the relative effect of each method on the performance of the calibration model. PLS method was used to regress the preprocessed data set and the SSCs of samples, and the cross-validation was to select the optimal number of PLS factors. Smoothing and scattering corection were essential in increasing the prediction performance of PLS regression model and the OSC contributed to reduction of the number of PLS factors. The first derivative resulted in unfavorable effect on the prediction performance. MSC and SNV showed similar effect. A robust calibration model could be developed by the preprocessing combination of Savitzky Golay smoothing, MSC and OSC, which resulted in SEP= 0.507, bias=0.032 and R$^2$=0.8823.

  • PDF

동일 데이터를 이용한 구조방정식(AMOS, LISREL and PLS) 툴 간의 비교분석 (A Comparison Analysis among Structural Equation Modeling (AMOS, LISREL and PLS) using the Same Data)

  • 남수태;김도관;진찬용
    • 한국정보통신학회:학술대회논문집
    • /
    • 한국정보통신학회 2018년도 춘계학술대회
    • /
    • pp.131-134
    • /
    • 2018
  • 구조방정식 모델링은 경로분석 및 확인적 요인분석을 동시에 수행해 주는 통계적 절차를 따르고 있다. 오늘날이 통계적 절차는 사회과학 분야의 연구자에게 필수적인 도구이다. 구조방정식 모델링 분석을 해주는 대표적인 도구로는 (AMOS, LISREL and PLS)가 있다. AMOS는 초보자가 사용할 수 있도록 편리한 그래픽 사용자 인터페이스를 제공해 주고 있다. PLS는 그래픽 사용자 인터페이스뿐만 아니라 정규분포에 대한 제약조건도 없다는 장점을 가지고 있다. 또한 사회과학 분야에서 가장 많이 사용하는 3가지 도구(Applications)를 비교분석 하였다. 이러한 결과를 바탕으로 연구의 한계와 시사점을 제시하고자 한다.

  • PDF

다변량 통계분석법을 이용한 PET 중합공정 중 직접 에스테르화 반응기의 거동 및 생산제품 예측 (Multivariate Statistical Analysis Approach to Predict the Reactor Properties and the Product Quality of a Direct Esterification Reactor for PET Synthesis)

  • 김성영;정창복;최수형;이범석;이범석
    • 제어로봇시스템학회논문지
    • /
    • 제11권6호
    • /
    • pp.550-557
    • /
    • 2005
  • The multivariate statistical analysis methods, using both multiple linear regression(MLR) and partial least square(PLS), have been applied to predict the reactor properties and the product quality of a direct esterification reactor for polyethylene terephthalate(PET) synthesis. On the basis of the set of data including the flow rate of water vapor, the flow rate of EG vapor, the concentration of acid end groups of a product and other operating conditions such as temperature, pressure, reaction times and feed monomer mole ratio, two multi-variable analysis methods have been applied. Their regression and prediction abilities also have been compared. The prediction results are critically compared with the actual plant data and the other mathematical model based results in reliability. This paper shows that PLS method approach can be used for the reasonably accurate prediction of a product quality of a direct esterification reactor in PET synthesis process.

정비예정구역 해제지역 재생사업의 정비요소와 고령거주자의 사업 만족도 간의 영향관계 사례연구 - 서울시 연남동, 북가좌동 시범사업지를 중심으로 - (A Case Study of Housing Regeneration Projects in Yonnam-dong and Buk Gajwa-dong, Seoul: The Determinants of Satisfaction of Elderly Residents)

  • 김아름;구자훈
    • 한국주거학회논문집
    • /
    • 제27권5호
    • /
    • pp.11-23
    • /
    • 2016
  • The purpose of this study is to establish the determinants of satisfaction with the results of housing regeneration projects among their elderly residents, and to suggest the political implications. The survey included questionnaires about satisfaction levels with the projects' physical and non-physical maintenance factors. The results were statistically analyzed by correlation analysis and PLS regression analysis. As a result of the study, firstly, the physical factors rather than non-physical factors (such as home improvement and management support, community support, the economic foundations and professional support) were found to have a large effect on elderly residents' satisfaction. Secondly, the non-physical factors, such as economic factors were analyzed among senior job offers that are both highly influential in the two regions Yonnam-dong and Bukgajwa-dong. Finally, electrical maintenance work, tree planting, a "Green" parking plan, or refuse the effect of visually larger landscape improvement, such as bins installed, maintenance of local factors that contribute to the greenery of the area were judged to be important.

Selecting Significant Wavelengths to Predict Chlorophyll Content of Grafted Cucumber Seedlings Using Hyperspectral Images

  • Jang, Sung Hyuk;Hwang, Yong Kee;Lee, Ho Jun;Lee, Jae Su;Kim, Yong Hyeon
    • 대한원격탐사학회지
    • /
    • 제34권4호
    • /
    • pp.681-692
    • /
    • 2018
  • This study was performed to select the significant wavelengths for predicting the chlorophyll content of grafted cucumber seedlings using hyperspectral images. The visible and near-infrared (VNIR) images and the short-wave infrared images of cucumber cotyledon samples were measured by two hyperspectral cameras. A correlation coefficient spectrum (CCS), a stepwise multiple linear regression (SMLR), and partial least squares (PLS) regression were used to determine significant wavelengths. Some wavelengths at 501, 505, 510, 543, 548, 619, 718, 723, and 727 nm were selected by CCS, SMLR, and PLS as significant wavelengths for estimating chlorophyll content. The results from the calibration models built by SMLR and PLS showed fair relationship between measured and predicted chlorophyll concentration. It was concluded that the hyperspectral imaging technique in the VNIR region is suggested effective for estimating the chlorophyll content of grafted cucumber leaves, non-destructively.