Search | Korea Science

Learning Text Chunking Using Maximum Entropy Models (최대 엔트로피 모델을 이용한 텍스트 단위화 학습)

Park, Seong-Bae;Zhang, Byoung-Tak
- Annual Conference on Human and Language Technology
- /
- 2001.10d
- /
- pp.130-137
- /
- 2001
최대 엔트로피 모델(maximum entropy model)은 여러 가지 자연언어 문제를 학습하는데 성공적으로 적용되어 왔지만, 두 가지의 주요한 문제점을 가지고 있다. 그 첫번째 문제는 해당 언어에 대한 많은 사전 지식(prior knowledge)이 필요하다는 것이고, 두번째 문제는 계산량이 너무 많다는 것이다. 본 논문에서는 텍스트 단위화(text chunking)에 최대 엔트로피 모델을 적용하는 데 나타나는 이 문제점들을 해소하기 위해 새로운 방법을 제시한다. 사전 지식으로, 간단한 언어 모델로부터 쉽게 생성된 결정트리(decision tree)에서 자동적으로 만들어진 규칙을 사용한다. 따라서, 제시된 방법에서의 최대 엔트로피 모델은 결정트리를 보강하는 방법으로 간주될 수 있다. 계산론적 복잡도를 줄이기 위해서, 최대 엔트로피 모델을 학습할 때 일종의 능동 학습(active learning) 방법을 사용한다. 전체 학습 데이터가 아닌 일부분만을 사용함으로써 계산 비용은 크게 줄어 들 수 있다. 실험 결과, 제시된 방법으로 결정트리의 오류의 수가 반으로 줄었다. 대부분의 자연언어 데이터가 매우 불균형을 이루므로, 학습된 모델을 부스팅(boosting)으로 강화할 수 있다. 부스팅을 한 후 제시된 방법은 전문가에 의해 선택된 자질로 학습된 최대 엔트로피 모델보다 졸은 성능을 보이며 지금까지 보고된 기계 학습 알고리즘 중 가장 성능이 좋은 방법과 비슷한 성능을 보인다 텍스트 단위화가 일반적으로 전체 구문분석의 전 단계이고 이 단계에서의 오류가 다음 단계에서 복구될 수 없으므로 이 성능은 텍스트 단위화에서 매우 의미가 길다.
PDF

A Study on Development of the Dual-thrust Flight Motor for Enhancing the Hit Probability (명중률 향상을 위한 이중추력형 비행모터 개발에 대한 연구)

Kim, Hanjun;Kim, Eunmi;Kim, Namsik;Lee, Wonbok;Yang, Youngjun
- Journal of the Korean Society of Propulsion Engineers
- /
- v.18 no.4
- /
- pp.74-80
- /
- 2014
This paper describes the development of the dual-thrust flight motor for enhancing the hit probability of unguided rockets. We designed dual-thrust flight motor by shape modification of the double base propellant with high burning rate, and confirmed the dual-thrust performance by static firing tests. The test results showed the thrust ratio of about 1:7.6 between sustaining phase and boosting phase, and had a quietly normal dual-thrust characteristics. And the results showed that there was not the fire extinction phenomenon of propellant due to the pressure drop.
https://doi.org/10.6108/KSPE.2014.18.4.074 인용 PDF KSCI

Context-Aware Fusion with Support Vector Machine (Support Vector Machine을 이용한 문맥 인지형 융합)

Heo, Gyeong-Yong;Kim, Seong-Hoon
- Journal of the Korea Society of Computer and Information
- /
- v.19 no.6
- /
- pp.19-26
- /
- 2014
An ensemble classifier system is a widely-used multi-classifier system, which combines the results from each classifier and, as a result, achieves better classification result than any single classifier used. Several methods have been used to build an ensemble classifier including boosting, which is a cascade method where misclassified examples in previous stage are used to boost the performance in current stage. Boosting is, however, a serial method which does not form a complete feedback loop. In this paper, proposed is context sensitive SVM ensemble (CASE) which adopts SVM, one of the best classifiers in term of classification rate, as a basic classifier and clustering method to divide feature space into contexts. As CASE divides feature space and trains SVMs simultaneously, the result from one component can be applied to the other and CASE achieves better result than boosting. Experimental results prove the usefulness of the proposed method.
https://doi.org/10.9708/jksci.2014.19.6.019 인용 PDF KSCI

Performance Analysis of Trading Strategy using Gradient Boosting Machine Learning and Genetic Algorithm

Jang, Phil-Sik
- Journal of the Korea Society of Computer and Information
- /
- v.27 no.11
- /
- pp.147-155
- /
- 2022
In this study, we developed a system to dynamically balance a daily stock portfolio and performed trading simulations using gradient boosting and genetic algorithms. We collected various stock market data from stocks listed on the KOSPI and KOSDAQ markets, including investor-specific transaction data. Subsequently, we indexed the data as a preprocessing step, and used feature engineering to modify and generate variables for training. First, we experimentally compared the performance of three popular gradient boosting algorithms in terms of accuracy, precision, recall, and F1-score, including XGBoost, LightGBM, and CatBoost. Based on the results, in a second experiment, we used a LightGBM model trained on the collected data along with genetic algorithms to predict and select stocks with a high daily probability of profit. We also conducted simulations of trading during the period of the testing data to analyze the performance of the proposed approach compared with the KOSPI and KOSDAQ indices in terms of the CAGR (Compound Annual Growth Rate), MDD (Maximum Draw Down), Sharpe ratio, and volatility. The results showed that the proposed strategies outperformed those employed by the Korean stock market in terms of all performance metrics. Moreover, our proposed LightGBM model with a genetic algorithm exhibited competitive performance in predicting stock price movements.
https://doi.org/10.9708/jksci.2022.27.11.147 인용 PDF KSCI HTML

A Sign Language Translator using Data Mining in Kinect Environment (키넥트 환경에서 데이터 마이닝을 이용한 수화 번역기)

Lee, Sang-Jun;Woo, Tea-Ho;Kim, Jia;Park, Seon-Yeong;Lee, Soo-Won;Kim, Gye-Young
- Proceedings of the Korea Information Processing Society Conference
- /
- 2012.11a
- /
- pp.619-622
- /
- 2012
본 연구에서는 키넥트(Kinect) 센서를 통해 수화 동작에서 손의 좌표와 이동방향을 추출하여 속성으로 하고, 데이터 마이닝의 분류 기법을 통해 수화를 인식하여 그 결과를 한글 텍스트로 번역해주는 소프트웨어를 개발한다. 제안 방법의 1단계에서는 0.05초 단위로 추출한 손의 좌표만을 속성으로 한다. 2단계에서는 개개인의 특성 및 화면상의 위치와 같은 요소에 따라 좌표 값이 달라지기 때문에, 손의 움직임에서 변위를 추출하여 손이 움직이는 방향을 속성으로 한다. 하지만 비슷한 방향으로 움직이는 수화가 있을 경우 수화의 구분이 어려우므로 3단계에서는 손의 좌표, 방향 두 가지를 분류하는 속성으로 사용한다. 향후 연구 방향은 수화의 중요한 요소인 손의 위치를 속성으로 추가시키고, 데이터 마이닝의 부스팅(Boosting) 기법을 적용하여 인식률을 높이는 것이다.
https://doi.org/10.3745/PKIPS.y2012m11a.619 인용 PDF

Boosted DNA Computing for Evolutionary Graphical Structure Learning (진화하는 그래프 구조 학습을 위한 부스티드 DNA 컴퓨팅)

Seok Ho-Sik;Zhang Byoung-Tak
- Proceedings of the Korean Information Science Society Conference
- /
- 2005.07b
- /
- pp.265-267
- /
- 2005
DNA 컴퓨팅은 분자 수준(molecular level)에서 연산을 수행한다. 따라서 일반적인 실리콘 기반의 컴퓨터에서와는 달리, 순차적인 연산 제어를 보장하기 어렵다는 특징이 있다. 그러나 DNA 컴퓨팅은 화학반응에 기초한 연산이기 때문에, 실험자가 의도한 연산을 많은 수의 분자에 동시에 적용할 수 있으므로 실리콘 기반의 컴퓨터와는 비교할 수 없는 병렬 연산을 구현할 수 있다. 병렬 연산을 구현하고자 할 때, 일반적으로 연산에 사용하는 모든 DNA 분자들을 대상으로 연산을 구현할 수도 있다. 그러나 전체가 아닌 일부의 분자들을 상대로 연산을 수행하는 것 역시 가능하며 이 때 자연스러운 방법으로 사용할 수 있는 방법이 배깅(Bagging)이나 부스팅(Boosting)과 같은 앙상블(ensemble) 계열의 학습 방법이다. 일반적인 부스팅과 달리 가중치를 부여하는 것이 아니라 특정 학습자(learner)를 나타내는 분자들을 증폭한다면 가중치를 분자의 양으로 표현하는 것이 가능하므로 분자 수준에서 앙상블 계열의 학습을 구현하는 것이 가능하다. 본 논문에서는 앙상블 계열의 학습 방법 중 특히 부스팅의 효과를 DNA 컴퓨팅에 응용하고자 할 때, 어떤 방법이 가능하며, 표현 과정에서 고려해야 할 사항은 어떠한 것들이 있는지 고려하고자 한다. 본 논문에서는 규모를 사전에 한정할 수 없는 진화 가능한 그래프 구조(evolutionary graph structure)를 학습할 수 있는 방법을 찾아보고자 한다. 진화 가능한 그래프 구조는 기존의 DNA 컴퓨팅 방법으로는 학습할 수 없는 문제이다. 그러나 조합 가능한 수를 사전에 정의할 수 없기 때문에 분자의 수에 상관없이 동일한 연산 시간에 문제를 해결할 수 있는 DNA 컴퓨팅의 장정을 가장 잘 발휘할 수 있는 문제이기도 하다.개별 태스크의 특성에 따른 성능 조절과 태스크의 변화에 따른 빠른 반응을 자랑으로 한다. 본 논문에선 TIB 알고리즘을 리눅스 커널에 구현하여 성능을 평가하였고 그 결과 리눅스에서 사용되는 기존 인터벌 기반의 알고리즘들에 비해 좋은 전력 절감 효과를 얻을 수 있었다.과는 한식 외식업체들이 고객들의 재구매 의도를 높이기 위해서는 한식 외식업체의 서비스요인, 식음료요인, 이벤트 요인 등을 강화함으로써 전반적인 종사원 서비스 품질과 식음료품질을 높이는 전략을 취해야 한다는 것을 시사해주고 있다. 본 연구는 대구 경북소재 한식 외식업체만을 대상으로 하여 연구를 실시하여 연구의 일반화와 한식 외식업체를 이용하는 이용 고객들이 한식 외식업체를 재방문하는 재구매 의도가 발생하는데 있어 발생하는 과정을 설명하는 종단적 연구를 실시하지 못한 한계점을 가지고 있다.아직 산업 디자인이 품질경쟁력에 크게 영향을 미치는 성숙단계에 이르지 못하였음을 의미한다. (2) 제품 디자인에게 영향을 끼치는 유의적인 변수는 연구개발력, 연구개발투자 수준, 혁신활동 수준(5S, TPM, 6Sigma 운동, QC 등)이며, 제품 디자인은 우선 품질경쟁력을 높여 간접적으로 고객만족과 고객 충성을 유발하는 것으로 추정되었다. 상기의 분석결과로부터, 본 연구는 다음과 같은 정책적 함의를 도출하였다. 첫째, 신상품 개발과 혁신을 위한 포괄적인 연구개발 프로젝트를 품질 경쟁력의 주요 결정요인(제품의 기본성능, 신뢰성, 수명(내구성) 및 제품 디자인)과 연계하여 추진해야 할 것이다. 둘째, 기업은 디자인 경영 마인드 제고와 디자인 전문인력 양성을, 대학은 디자인 현장 업무를 통하여 창의력 증진과 기획 및 마케팅 능력 교육을, 정부는 디자인 기술개발 및 디자인 교육지원의 강화를 통하여 각각 디자인 경쟁력$\righta
PDF

Prediction of Cryptocurrency Price Trend Using Gradient Boosting (그래디언트 부스팅을 활용한 암호화폐 가격동향 예측)

Heo, Joo-Seong;Kwon, Do-Hyung;Kim, Ju-Bong;Han, Youn-Hee;An, Chae-Hun
- KIPS Transactions on Software and Data Engineering
- /
- v.7 no.10
- /
- pp.387-396
- /
- 2018
Stock price prediction has been a difficult problem to solve. There have been many studies to predict stock price scientifically, but it is still impossible to predict the exact price. Recently, a variety of types of cryptocurrency has been developed, beginning with Bitcoin, which is technically implemented as the concept of distributed ledger. Various approaches have been attempted to predict the price of cryptocurrency. Especially, it is various from attempts to stock prediction techniques in traditional stock market, to attempts to apply deep learning and reinforcement learning. Since the market for cryptocurrency has many new features that are not present in the existing traditional stock market, there is a growing demand for new analytical techniques suitable for the cryptocurrency market. In this study, we first collect and process seven cryptocurrency price data through Bithumb's API. Then, we use the gradient boosting model, which is a data-driven learning based machine learning model, and let the model learn the price data change of cryptocurrency. We also find the most optimal model parameters in the verification step, and finally evaluate the prediction performance of the cryptocurrency price trends.
https://doi.org/10.3745/KTSDE.2018.7.10.387 인용 PDF KSCI

Development of the Dual-Thrust Rocket Motor (이중추력형 로켓 모터의 개발)

Lee, Do-Hyung;Yoon, Myong-Won;Hwang, Kab-Sung
- Journal of the Korean Society for Aeronautical & Space Sciences
- /
- v.32 no.9
- /
- pp.130-135
- /
- 2004
This paper describes the development of the dual-thrust rocket motor, which gets a significant change in thrust level by varying the burning area of the propellant grain. Reduced smoke propellant of low burning rate was formulated and the finocyl type grain was designed to get the boosting- and sustaining-phase of the thrust level. And the motor firing data were analyzed in detail. Developed motor was applied to the missile system to implement the successful flight test and this development helped to upgrade the performance of the missile system. The results will be usefully applied to the development of the similar rocket motors.
https://doi.org/10.5139/JKSAS.2004.32.9.130 인용 PDF KSCI

A study on the performance prediction technique of the dual-thrust rocket motor (이중 추력형 로켓모타의 성능예측 기법 연구)

이도형
- Journal of the Korean Society of Propulsion Engineers
- /
- v.5 no.2
- /
- pp.38-43
- /
- 2001
In this study, the technique of the performance prediction on the finocyl-type dual-thrust rocket motor is developed, and the predicted data are compared with those of the static firing tests. The prediction is carried out with the separate calculations of the grain burning area and the performance of the rocket motor. When predicting the performance of the dual-thrust rocket motor, the different correction factors should be used at the boosting and sustaining phases. Otherwise, an error of prediction will follow. Reprediction using the separate correction factors shows good agreement with the test data within 0.5% error.
PDF

Prediction of golf scores on the PGA tour using statistical models (PGA 투어의 골프 스코어 예측 및 분석)

Lim, Jungeun;Lim, Youngin;Song, Jongwoo
- The Korean Journal of Applied Statistics
- /
- v.30 no.1
- /
- pp.41-55
- /
- 2017
This study predicts the average scores of top 150 PGA golf players on 132 PGA Tour tournaments (2013-2015) using data mining techniques and statistical analysis. This study also aims to predict the Top 10 and Top 25 best players in 4 different playoffs. Linear and nonlinear regression methods were used to predict average scores. Stepwise regression, all best subset, LASSO, ridge regression and principal component regression were used for the linear regression method. Tree, bagging, gradient boosting, neural network, random forests and KNN were used for nonlinear regression method. We found that the average score increases as fairway firmness or green height or average maximum wind speed increases. We also found that the average score decreases as the number of one-putts or scrambling variable or longest driving distance increases. All 11 different models have low prediction error when predicting the average scores of PGA Tournaments in 2015 which is not included in the training set. However, the performances of Bagging and Random Forest models are the best among all models and these two models have the highest prediction accuracy when predicting the Top 10 and Top 25 best players in 4 different playoffs.
https://doi.org/10.5351/KJAS.2017.30.1.041 인용 PDF KSCI

Search Result 13, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)