통합 검색 | Korea Science

Compositional data analysis by the square-root transformation: Application to NBA USG% data

Jeseok Lee;Byungwon Kim
- Communications for Statistical Applications and Methods
- /
- 제31권3호
- /
- pp.349-363
- /
- 2024
Compositional data refers to data where the sum of the values of the components is a constant, hence the sample space is defined as a simplex making it impossible to apply statistical methods developed in the usual Euclidean vector space. A natural approach to overcome this restriction is to consider an appropriate transformation which moves the sample space onto the Euclidean space, and log-ratio typed transformations, such as the additive log-ratio (ALR), the centered log-ratio (CLR) and the isometric log-ratio (ILR) transformations, have been mostly conducted. However, in scenarios with sparsity, where certain components take on exact zero values, these log-ratio type transformations may not be effective. In this work, we mainly suggest an alternative transformation, that is the square-root transformation which moves the original sample space onto the directional space. We compare the square-root transformation with the log-ratio typed transformation by the simulation study and the real data example. In the real data example, we applied both types of transformations to the USG% data obtained from NBA, and used a density based clustering method, DBSCAN (density-based spatial clustering of applications with noise), to show the result.
https://doi.org/10.29220/CSAM.2024.31.3.349 인용 PDF

Effect of zero imputation methods for log-transformation of independent variables in logistic regression

Seo Young Park
- Communications for Statistical Applications and Methods
- /
- 제31권4호
- /
- pp.409-425
- /
- 2024
Logistic regression models are commonly used to explain binary health outcome variable using independent variables such as patient characteristics in medical science and public health research. Although there is no distributional assumption required for independent variables in logistic regression, variables with severely right-skewed distribution such as lab values are often log-transformed to achieve symmetry or approximate normality. However, lab values often have zeros due to limit of detection which makes it impossible to apply log-transformation. Therefore, preprocessing to handle zeros in the observation before log-transformation is necessary. In this study, five methods that remove zeros (shift by 1, shift by half of the smallest nonzero, shift by square root of the smallest nonzero, replace zeros with half of the smallest nonzero, replace zeros with the square root of the smallest nonzero) are investigated in logistic regression setting. To evaluate performances of these methods, we performed a simulation study based on randomly generated data from log-normal distribution and logistic regression model. Shift by 1 method has the worst performance, and overall shift by half of the smallest nonzero method, replace zeros with half of the smallest nonzero method, and replace zeros with the square root of the smallest nonzero method showed comparable and stable performances.
https://doi.org/10.29220/CSAM.2024.31.4.409 인용 PDF

ARMA(p, q) 모형에서 멱변환의 재변환에 관한 연구 - 모의실험을 중심으로 (Re-Transformation of Power Transformation for ARMA(p, q) Model - Simulation Study)

강전훈;신기일
- 응용통계연구
- /
- 제28권3호
- /
- pp.511-527
- /
- 2015
ARMA(p, q) 모형 분석에서 분산 안정화 또는 정규화를 위해 멱변환(power transformation)이 사용된다. 변환된 자료를 이용하여 분석이 이루어지며 원 자료의 예측을 위해 재변환이 사용된다. 이때 흔히 변환된 자료 분석에서 얻어진 예측값의 역함수 값이 원자료 예측값으로 사용되지만 이는 편향이 있는 것으로 알려져 있다. 이를 해결하기 위해 로그 변환의 경우 Granger과 Newbold (1976)는 로그-정규분포의 기댓값을 이용할 것을 제안하였다. 본 연구에서는 모의실험을 통하여 제곱근 변환과 로그 변환 후 재변환을 사용할 때 예측값으로 기댓값의 역함수를 이용하는 방법과 역함수의 기댓값을 사용하였을 때의 추정의 결과를 모의실험을 통하여 비교하였다.
https://doi.org/10.5351/KJAS.2015.28.3.511 인용 PDF KSCI

로그폴라 사상과 어파인 변환을 이용한 새로운 템플릿 기반 얼굴 인식 (New Template Based Face Recognition Using Log-polar Mapping and Affine Transformation)

김문갑;최일;진성일
- 대한전자공학회논문지SP
- /
- 제39권2호
- /
- pp.1-10
- /
- 2002
이 논문에서는 크기와 영상 평면상에서 회전 (in-plane rotation) 변화를 가지는 정면 얼굴 영상의 인식성능을 향상시키기 위하여, 새로운 템플릿 (template) 기반 접근 방법들을 제안한다. 인식 성능을 향상시키기 위한 템플릿들은 크기와 회전 변화가 다른 다수의 영상들을 선형 또는 비선형 연산에 의하여 생성된다. 얼굴의 크기와 영상 평면에서 회전 변화에 무관한 얼굴의 특징을 추출하기 위하여 어파인 (affine) 변환, 로그폴라 (log-polar) 사상, 그리고 로그폴라 영상에 기반한 FFT들이 이용된다. 제안된 방법들은 인식률과 수행 시간 측면에서 비교된다. 실험 결과로부터 제안된 템플릿을 이용한 방법들의 인식률이 한 장의 영상으로 생성된 템플릿을 이용한 방법들의 인식률보다 우수함을 나타낸다. 어파인 변환을 이용한 방법의 인식률이 로그폴라 사상을 이용한 방법과 로그폴라 영상에 기반한 FFT 방법의 인식률보다 우수하며, 수행 시간 측면에서는 로그폴라 사상을 이용한 방법이 가장 빠르다.
PDF KSCI

로그변환 모델에 따른 생물학적 동등성 판정 연구 (Analysis of Bioequivalence Study using a Log-transformed Model)

이영주;김윤균;이명걸;정석재;이민화;심창구
- 약학회지
- /
- 제44권4호
- /
- pp.308-314
- /
- 2000
Logarithmic transformation of pharmacokinetic parameters is routinely used in bioequivalence studies based on pharmacokinetic and statistical grounds by the United States Food and Drug Administration (FDA), European Committee for Proprietary Medicinal Products (CPMP), and Japanese National Institute of Health and Science (NIHS). Although it has not yet been recommended by the Korea Food and Drug Administration (KFDA), its use is becoming increasingly necessary in order to harmonize with international standards. In the present study, statistical procedures for the analysis of a bioequivalence based on the log transformation and a related SAS procedure were demonstrated in order to aid the understanding and application. The AUC parameters used in this demonstration were taken from the previous bioequivalence study for two aceclofenac tablets, which were performed in a single-dose crossover design. Analysis of variance (ANOVA), statistical power to detect 20% difference between the tablets, minimum detectable difference and confidence intervals were all assessed following log-transformation of the data. Bioequivalence of two aceclofenac tablets was then estimated based on the guideline of FDA. Considering the international effort for harmaonization of guidelines for bioequivalence tests, this approach may require a further evaluation for a future adaptation in the Korea Guidelines of Bioequivalence Tests (KGBT).
PDF

변수변환을 통한 포항지역 미세먼지의 통계적 예보모형에 관한 연구 (A Study on Statistical Forecasting Models of PM10 in Pohang Region by the Variable Transformation)

이영섭;김현구;박종석;김희경
- 한국대기환경학회지
- /
- 제22권5호
- /
- pp.614-626
- /
- 2006
Using the data of three environmental monitoring sites in Pohang area(KME112, KME113, and KME114), statistical forecasting models of the daily maximum and mean values of PM10 have been developed. Since the distributions of the daily maximum and mean PM10 values are skewed, which are similar to the Weibull distribution, these values were log-transformed to increase prediction accuracy by approximating the normal distribution. Three statistical forecasting models, which are regression, neural networks(NN) and support vector regression(SVR), were built using the log-transformed response variables, i.e., log(max(PM10)) or log(mean (PM10)). Also, the forecasting models were validated by the measure of RMSE, CORR, and IOA for the model comparison and accuracy. The improvement rate of IOA before and after the log-transformation in the daily maximum PM10 prediction was 12.7% for the regression and 22.5% for NN. In particular, 42.7% was improved for SVR method. In the case of the daily mean PM10 prediction, IOA value was improved by 5.1% for regression, 6.5% for NN, and 6.3% for SVR method. As a conclusion, SVR method was found to be performed better than the other methods in the point of the model accuracy and fitness views.
PDF KSCI

빈도해석에 의한 용담 수위관측소 지점의 갈수량 분석 (Frequency Analysis of Low Flows at Yongdam Stage Station)

안태진;여운식;정광근
- 한국관개배수논문집
- /
- 제5권1호
- /
- pp.20-31
- /
- 1998
The Power transformation, the modified Power transformation, the logarithmic transformation, the square-root logarithmic transformation, the SMEMAX transformation, the Extreme value type III, the Weibull, the log Pearson type III, the lognormal distributi
PDF

수생태 독성자료의 정규성 분포 특성 확인을 통해 통계분석 시 분포 특성 적용에 대한 타당성 확인 연구 (The Validation Study of Normality Distribution of Aquatic Toxicity Data for Statistical Analysis)

옥승엽;문효방;나진성
- 한국환경보건학회지
- /
- 제45권2호
- /
- pp.192-202
- /
- 2019
Objectives: According to the central limit theorem, the samples in population might be considered to follow normal distribution if a large number of samples are available. Once we assume that toxicity dataset follow normal distribution, we can treat and process data statistically to calculate genus or species mean value with standard deviation. However, little is known and only limited studies are conducted to investigate whether toxicity dataset follows normal distribution or not. Therefore, the purpose of study is to evaluate the generally accepted normality hypothesis of aquatic toxicity dataset Methods: We selected the 8 chemicals, which consist of 4 organic and 4 inorganic chemical compounds considering data availability for the development of species sensitivity distribution. Toxicity data were collected at the US EPA ECOTOX Knowledgebase by simple search with target chemicals. Toxicity data were re-arranged to a proper format based on the endpoint and test duration, where we conducted normality test according to the Shapiro-Wilk test. Also we investigated the degree of normality by simple log transformation of toxicity data Results: Despite of the central limit theorem, only one large dataset (n>25) follow normal distribution out of 25 large dataset. By log transforming, more 7 large dataset show normality. As a result of normality test on small dataset (n<25), log transformation of toxicity value generally increases normality. Both organic and inorganic chemicals show normality growth for 26 species and 30 species, respectively. Those 56 species shows normality growth by log transformation in the taxonomic groups such as amphibian (1), crustacean (21), fish (22), insect (5), rotifer (2), and worm (5). In contrast, mollusca shows normality decrease at 1 species out of 23 that originally show normality. Conclusions: The normality of large toxicity dataset was not always satisfactory to the central limit theorem. Normality of those data could be improved through log transformation. Therefore, care should be taken when using toxicity data to induce, for example, mean value for risk assessment.
https://doi.org/10.5668/JEHS.2019.45.2.192 인용 PDF KSCI

Effects of Spectral Transformations on Leaf C:N Ratio Inversion with Hyperspectral Data

Run-he, SHI;Da-fang, ZHUANG;Qiao-jing, QIAN;Zheng, NIU
- 대한원격탐사학회:학술대회논문집
- /
- 대한원격탐사학회 2003년도 Proceedings of ACRS 2003 ISRS
- /
- pp.322-324
- /
- 2003
Leaf C:N ratio is a new factor in the field of biochemical inversion with hyperspectral data. Effects of common-used spectral transformations including log(R), log(1/R), 1/R, etc. from 400nm to 2490nm on its inversion are compared. Results show that their effects on statistical modeling are not apparent. Continuum removal is used on original reflectance in the range of 2030nm to 2220nm, in which exists an apparent absorption peak due to cellulose, lignin, protein, etc. The effect is distinctive and tends to improve the precision of C:N ratio inversion. Further, it is a robust and physically based transformation.
PDF

$n^3$ 프로세서 재구성가능 메쉬에서 $n^2$ 화소 이진영상과 경계코드간의 효율적인 변환 (Efficient Transformations Between an $n^2$ Pixel Binary Image and a Boundary Code on an $n^3$ Processor Reconfigurable Mesh)

김명
- 한국정보처리학회논문지
- /
- 제5권8호
- /
- pp.2027-2040
- /
- 1998
본 논문에서는 $n\timesn\timesn$ 프로세서로 구성된 재구성가능 메쉬에서 $n\timesn$개의 화소가 있는 이진영상을 경계코드로 변환 하거나 그 역변환을 하는 알고리즘을 제안한다. 이와 동일한 변환을 하는 O(1) 시간 알고리즘들이 이미 제안되었는데, 이들이 사용하는 프로세서의 수는 $O(n^4)$으로, 영상의 화소 수와 비교해 볼 때 지나치게 많다고 하겠다. 본 논문에서는 $n^3개의 프로세서만을 사용하는 속도 빠른 변환 알고리즘을 소개한다. 여기서 제안하는 경계코드를 이진영상으로 변환하는 알고리즘의 실행시간은 O(1)이고, 그 역변환 알고리즘의 실행시간은 O(log n)이다.
PDF

검색결과 96건 처리시간 0.033초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)