• Title/Summary/Keyword: regression using

검색결과 18,787건 처리시간 0.042초

잭나이프 및 붓스트랩 방법을 이용한 임상자료의 회귀계수 타당성 확인 (Check for regression coefficient using jackknife and bootstrap methods in clinical data)

  • 손기철;신임희
    • Journal of the Korean Data and Information Science Society
    • /
    • 제23권4호
    • /
    • pp.643-648
    • /
    • 2012
  • 여러 임상자료를 이용하여 반응변수와 설명변수간의 관계를 규명하는 분석이 많이 이루어지고 있다. 이를 위해서 회귀분석이 흔히 사용되고 있으며, 이를 통해 설명변수가 반응변수를 얼마나 설명하는지 또한 모형이 얼마나 자료에 적합한지에 대해 분석하고 있다. 그러나 임상자료로 분석된 회귀모형에 대한 타당성 확인은 대부분 분석된 회귀모형이 얼마나 자료를 설명하는가를 나타내는 결정계수만을 살펴보는 것에 그치고 있다. 결정계수 이외의 다른 방법으로도 분석된 회귀모형의 회귀계수에 대한 타당성을 확인할 필요가 있다. 따라서 본 논문에서는 잭나이프 회귀분석과 붓스트랩 회귀분석을 이용하여 임상자료로 분석한 회귀모형의 회귀계수에 대한 타당성을 확인하는 방법을 소개하고자 한다.

다중선형회귀법을 활용한 예민화와 환경변수에 따른 AL-6XN강의 공식특성 예측 (Prediction of Pitting Corrosion Characteristics of AL-6XN Steel with Sensitization and Environmental Variables Using Multiple Linear Regression Method)

  • 정광후;김성종
    • Corrosion Science and Technology
    • /
    • 제19권6호
    • /
    • pp.302-309
    • /
    • 2020
  • This study aimed to predict the pitting corrosion characteristics of AL-6XN super-austenitic steel using multiple linear regression. The variables used in the model are degree of sensitization, temperature, and pH. Experiments were designed and cyclic polarization curve tests were conducted accordingly. The data obtained from the cyclic polarization curve tests were used as training data for the multiple linear regression model. The significance of each factor in the response (critical pitting potential, repassivation potential) was analyzed. The multiple linear regression model was validated using experimental conditions that were not included in the training data. As a result, the degree of sensitization showed a greater effect than the other variables. Multiple linear regression showed poor performance for prediction of repassivation potential. On the other hand, the model showed a considerable degree of predictive performance for critical pitting potential. The coefficient of determination (R2) was 0.7745. The possibility for pitting potential prediction was confirmed using multiple linear regression.

다항식 회귀분석을 이용한 전자저울의 비선형 특성 개선 연구 (A Study of the Nonlinear Characteristics Improvement for a Electronic Scale using Multiple Regression Analysis)

  • 채규수
    • 융합정보논문지
    • /
    • 제9권6호
    • /
    • pp.1-6
    • /
    • 2019
  • 본 연구에서는 다항식 회귀분석(Polynomial regression analysis) 방법을 이용하여 비선형 특성을 갖는 전자저울의 질량 추정 모델 개발이 이루어 졌다. 전자저울에 사용되는 로드셀의 출력 단자 전압을 기준 질량 추를 사용하여 직접 측정하였고 이 데이터를 이용하여 MS Office 엑셀의 행렬식 계산과 데이터 추세선 분석 기능을 이용하여 다항식 회귀모델을 구하였다. 5kg까지 측정 가능한 로드셀 전자저울을 사용하여 100g단위로 질량을 측정하였고 다항식 회귀분석(Multiple regression analysis) 모델을 구하였으며, 단순(1차), 2차, 3차 다항식 회귀분석에 대한 오차를 구하였다. 각 모델에 대한 회귀 방정식의 적합도 분석을 위해 결정계수(Coefficient of determination)를 제시하여 추정 질량과 측정 데이터와의 상관관계를 나타내었다. 본 연구에서 제안하는 3차 다항식 모델을 이용하여 추정 값의 표준편차가 10g, 결정계수 1.0으로 상당히 정확한 모델을 얻었다. 본 연구에 사용된 선형 회귀 분석 이론을 바탕으로 최근 인공지능 분야에서 많이 사용되고 있는 로지스틱 회귀 분석(Logistic regression analysis)을 활용하여 기상예측, 신약개발, 경제지표 분석 등의 분야에 대한 다양한 연구를 수행할 수 있을 것으로 생각된다.

TIME SERIES PREDICTION USING INCREMENTAL REGRESSION

  • Kim, Sung-Hyun;Lee, Yong-Mi;Jin, Long;Chai, Duck-Jin;Ryu, Keun-Ho
    • 대한원격탐사학회:학술대회논문집
    • /
    • 대한원격탐사학회 2006년도 Proceedings of ISRS 2006 PORSEC Volume II
    • /
    • pp.635-638
    • /
    • 2006
  • Regression of conventional prediction techniques in data mining uses the model which is generated from the training step. This model is applied to new input data without any change. If this model is applied directly to time series, the rate of prediction accuracy will be decreased. This paper proposes an incremental regression for time series prediction like typhoon track prediction. This technique considers the characteristic of time series which may be changed over time. It is composed of two steps. The first step executes a fractional process for applying input data to the regression model. The second step updates the model by using its information as new data. Additionally, the model is maintained by only recent data in a queue. This approach has the following two advantages. It maintains the minimum information of the model by using a matrix, so space complexity is reduced. Moreover, it prevents the increment of error rate by updating the model over time. Accuracy rate of the proposed method is measured by RME(Relative Mean Error) and RMSE(Root Mean Square Error). The results of typhoon track prediction experiment are performed by the proposed technique IMLR(Incremental Multiple Linear Regression) is more efficient than those of MLR(Multiple Linear Regression) and SVR(Support Vector Regression).

  • PDF

MOISTURE CONTENT MEASUREMENT OF POWDERED FOOD USING RF IMPEDANCE SPECTROSCOPIC METHOD

  • Kim, K. B.;Lee, J. W.;S. H. Noh;Lee, S. S.
    • 한국농업기계학회:학술대회논문집
    • /
    • 한국농업기계학회 2000년도 THE THIRD INTERNATIONAL CONFERENCE ON AGRICULTURAL MACHINERY ENGINEERING. V.II
    • /
    • pp.188-195
    • /
    • 2000
  • This study was conducted to measure the moisture content of powdered food using RF impedance spectroscopic method. In frequency range of 1.0 to 30㎒, the impedance such as reactance and resistance of parallel plate type sample holder filled with wheat flour and red-pepper powder of which moisture content range were 5.93∼-17.07%w.b. and 10.87 ∼ 27.36%w.b., respectively, was characterized using by Q-meter (HP4342). The reactance was a better parameter than the resistance in estimating the moisture density defined as product of moisture content and bulk density which was used to eliminate the effect of bulk density on RF spectral data in this study. Multivariate data analyses such as principal component regression, partial least square regression and multiple linear regression were performed to develop one calibration model having moisture density and reactance spectral data as parameters for determination of moisture content of both wheat flour and red-pepper powder. The best regression model was one by the multiple linear regression model. Its performance for unknown data of powdered food was showed that the bias, standard error of prediction and determination coefficient are 0.179% moisture content, 1.679% moisture content and 0.8849, respectively.

  • PDF

Predictive analyses for balance and gait based on trunk performance using clinical scales in persons with stroke

  • Woo, Youngkeun
    • Physical Therapy Rehabilitation Science
    • /
    • 제7권1호
    • /
    • pp.29-34
    • /
    • 2018
  • Objective: This study aimed to predict balance and gait abilities with the Trunk Impairment scales (TIS) in persons with stroke. Design: Cross-sectional study. Methods: Sixty-eight participants with stoke were assessed with the TIS, Berg Balance scale (BBS), and Functional Gait Assessment (FGA) by a therapist. To describe of general characteristics, we used descriptive and frequency analyses, and the TIS was used as a predictive variable to determine the BBS. In the simple regression analysis, the TIS was used as a predictive variable for the BBS and FGA, and the TIS and BBS were used as predictive variables to determine the FGA in multiple regression analysis. Results: In the group with a BBS score of >45 for regression equation for predicting BBS score using TIS score, the coefficient of determination ($R^2$) was 0.234, and the $R^2$ was 0.500 in the group with a BBS score of ${\leq}45$. In the group with an FGA score >15 for regression equation for predicting FGA score using TIS score, the $R^2$ was 0.193, and regression equation for predicting FGA score using TIS score, the $R^2$ was 0.181 in the group of FGA score ${\leq}15$. In the group of FGA score >15 for regression equation for predicting FGA score using TIS and BBS score, the $R^2$ was 0.327. In the group of FGA score ${\leq}15$ for regression equation for predicting FGA score using TIS and BBS score, the $R^2$ was 0.316. Conclusions: The TIS scores are insufficient in predicting the FGA and BBS scores in those with higher balance ability, and the BBS and TIS could be used for predicting variables for FGA. However, TIS is a strong predictive variable for persons with stroke who have poor balance ability.

Inference of Gene Regulatory Networks via Boolean Networks Using Regression Coefficients

  • Kim, Ha-Seong;Choi, Ho-Sik;Lee, Jae-K.;Park, Tae-Sung
    • 한국생물정보학회:학술대회논문집
    • /
    • 한국생물정보시스템생물학회 2005년도 BIOINFO 2005
    • /
    • pp.339-343
    • /
    • 2005
  • Boolean networks(BN) construction is one of the commonly used methods for building gene networks from time series microarray data. However, BN has two major drawbacks. First, it requires heavy computing times. Second, the binary transformation of the microarray data may cause a loss of information. This paper propose two methods using liner regression to construct gene regulatory networks. The first proposed method uses regression based BN variable selection method, which reduces the computing time significantly in the BN construction. The second method is the regression based network method that can flexibly incorporate the interaction of the genes using continuous gene expression data. We construct the network structure from the simulated data to compare the computing times between Boolean networks and the proposed method. The regression based network method is evaluated using a microarray data of cell cycle in Caulobacter crescentus.

  • PDF

Imputation Method Using Local Linear Regression Based on Bidirectional k-nearest-components

  • Yonggeol, Lee
    • Journal of information and communication convergence engineering
    • /
    • 제21권1호
    • /
    • pp.62-67
    • /
    • 2023
  • This paper proposes an imputation method using a bidirectional k-nearest components search based local linear regression method. The bidirectional k-nearest-components search method selects components in the dynamic range from the missing points. Unlike the existing methods, which use a fixed-size window, the proposed method can flexibly select adjacent components in an imputation problem. The weight values assigned to the components around the missing points are calculated using local linear regression. The local linear regression method is free from the rank problem in a matrix of dependent variables. In addition, it can calculate the weight values that reflect the data flow in a specific environment, such as a blackout. The original missing values were estimated from a linear combination of the components and their weights. Finally, the estimated value imputes the missing values. In the experimental results, the proposed method outperformed the existing methods when the error between the original data and imputation data was measured using MAE and RMSE.

물류예측모형에 관한 연구 -수도권 물동량 예측을 중심으로- (A Study on Change of Logistics in the region of Seoul, Incheon, Kyunggi)

  • 노경호
    • 경영과정보연구
    • /
    • 제7권
    • /
    • pp.427-450
    • /
    • 2001
  • This research suggests the estimation methodology of Logistics. This paper elucidates the main problems associated with estimation in the regression model. We review the methods for estimating the parameters in the model and introduce a modified procedure in which all models are fitted and combined to construct a combination of estimates. The resulting estimators are found to be as efficient as the maximum likelihood (ML) estimators in various cases. Our method requires more computations but has an advantage for large data sets. Also, it enables to detect particular features in the data structure. Examples of real data are used to illustrate the properties of the estimators. The backgrounds of estimation of logistic regression model is the increasing logistic environment importance today. In the first phase, we conduct an exploratory study to discuss 9 independent variables. In the second phase, we try to find the fittest logistic regression model. In the third phase, we calculate the logistic estimation using logistic regression model. The parameters of logistic regression model were estimated using ordinary least squares regression. The standard assumptions of OLS estimation were tested. The calculated value of the F-statistics for the logistic regression model is significant at the 5% level. The logistic regression model also explains a significant amount of variance in the dependent variable. The parameter estimates of the logistic regression model with t-statistics in parentheses are presented in Table. The object of this paper is to find the best logistic regression model to estimate the comparative accurate logistics.

  • PDF

Local linear regression analysis for interval-valued data

  • Jang, Jungteak;Kang, Kee-Hoon
    • Communications for Statistical Applications and Methods
    • /
    • 제27권3호
    • /
    • pp.365-376
    • /
    • 2020
  • Interval-valued data, a type of symbolic data, is given as an interval in which the observation object is not a single value. It can also occur frequently in the process of aggregating large databases into a form that is easy to manage. Various regression methods for interval-valued data have been proposed relatively recently. In this paper, we introduce a nonparametric regression model using the kernel function and a nonlinear regression model for the interval-valued data. We also propose applying the local linear regression model, one of the nonparametric methods, to the interval-valued data. Simulations based on several distributions of the center point and the range are conducted using each of the methods presented in this paper. Various conditions confirm that the performance of the proposed local linear estimator is better than the others.