Search | Korea Science

Representative Labels Selection Technique for Document Cluster using WordNet (문서 클러스터를 위한 워드넷기반의 대표 레이블 선정 방법)

Kim, Tae-Hoon;Sohn, Mye
- Journal of Internet Computing and Services
- /
- v.18 no.2
- /
- pp.61-73
- /
- 2017
In this paper, we propose a Documents Cluster Labeling method using information content of words in clusters to understand what the clusters imply. To do so, we calculate the weight and frequency of the words. These two measures are used to determine the weight among the words in the cluster. As a nest step, we identify the candidate labels using the WordNet. At this time, the candidate labels are matched to least common hypernym of the words in the cluster. Finally, the representative labels are determined with respect to information content of the words and the weight of the words. To prove the superiority of our method, we perform the heuristic experiment using two kinds of measures, named the suitability of the candidate label ($Suitability_{cl}$) and the appropriacy of representative label ($Appropriacy_{rl}$). In applying the method proposed in this research, in case of suitability of the candidate label, it decreases slightly compared with existing methods, but the computational cost is about 20% of the conventional methods. And we confirmed that appropriacy of the representative label is better results than the existing methods. As a result, it is expected to help data analysts to interpret the document cluster easier.
https://doi.org/10.7472/jksii.2017.18.2.61 인용 PDF KSCI

A Study-on Context-Dependent Acoustic Models to Improve the Performance of the Korea Speech Recognition (한국어 음성인식 성능향상을 위한 문맥의존 음향모델에 관한 연구)

황철준;오세진;김범국;정호열;정현열
- Journal of the Institute of Convergence Signal Processing
- /
- v.2 no.4
- /
- pp.9-15
- /
- 2001
In this paper we investigate context dependent acoustic models to improve the performance of the Korean speech recognition . The algorithm are using the Korean phonological rules and decision tree, By Successive State Splitting(SSS) algorithm the Hidden Merkov Netwwork(HM-Net) which is an efficient representation of phoneme-context-dependent HMMs, can be generated automatically SSS is powerful technique to design topologies of tied-state HMMs but it doesn't treat unknown contexts in the training phoneme contexts environment adequately In addition it has some problem in the procedure of the contextual domain. In this paper we adopt a new state-clustering algorithm of SSS, called Phonetic Decision Tree-based SSS (PDT-SSS) which includes contexts splits based on the Korean phonological rules. This method combines advantages of both the decision tree clustering and SSS, and can generated highly accurate HM-Net that can express any contexts To verify the effectiveness of the adopted methods. the experiments are carried out using KLE 452 word database and YNU 200 sentence database. Through the Korean phoneme word and sentence recognition experiments. we proved that the new state-clustering algorithm produce better phoneme, word and continuous speech recognition accuracy than the conventional HMMs.
PDF

An Intelligent Marking System based on Semantic Kernel and Korean WordNet (의미커널과 한글 워드넷에 기반한 지능형 채점 시스템)

Cho Woojin;Oh Jungseok;Lee Jaeyoung;Kim Yu-Seop
- The KIPS Transactions:PartA
- /
- v.12A no.6 s.96
- /
- pp.539-546
- /
- 2005
Recently, as the number of Internet users are growing explosively, e-learning has been applied spread, as well as remote evaluation of intellectual capacity However, only the multiple choice and/or the objective tests have been applied to the e-learning, because of difficulty of natural language processing. For the intelligent marking of short-essay typed answer papers with rapidness and fairness, this work utilize heterogenous linguistic knowledges. Firstly, we construct the semantic kernel from un tagged corpus. Then the answer papers of students and instructors are transformed into the vector form. Finally, we evaluate the similarity between the papers by using the semantic kernel and decide whether the answer paper is correct or not, based on the similarity values. For the construction of the semantic kernel, we used latent semantic analysis based on the vector space model. Further we try to reduce the problem of information shortage, by integrating Korean Word Net. For the construction of the semantic kernel we collected 38,727 newspaper articles and extracted 75,175 indexed terms. In the experiment, about 0.894 correlation coefficient value, between the marking results from this system and the human instructors, was acquired.
https://doi.org/10.3745/KIPSTA.2005.12A.6.539 인용 PDF KSCI

Performance Improvement of Microphone Array Speech Recognition Using Features Weighted Mahalanobis Distance (가중특징 Mahalanobis거리를 이용한 마이크 어레이 음석인식의 성능향상)

Nguyen, Dinh Cuong;Chung, Hyun-Yeol
- The Journal of the Acoustical Society of Korea
- /
- v.29 no.1E
- /
- pp.45-53
- /
- 2010
In this paper, we present the use of the Features Weighted Mahalanobis Distance (FWMD) in improving the performance of Likelihood Maximizing Beamforming (Limabeam) algorithm in speech recognition for microphone array. The proposed approach is based on the replacement of the traditional distance measure in a Gaussian classifier with adding weight for different features in the Mahalanobis distance according to their distances after the variance normalization. By using Features Weighted Mahalanobis Distance for Limabeam algorithm (FWMD-Limabeam), we obtained correct word recognition rate of 90.26% for calibrate Limabeam and 87.23% for unsupervised Limabeam, resulting in a higher rate of 3% and 6% respectively than those produced by the original Limabearn. By implementing a HM-Net speech recognition strategy alternatively, we could save memory and reduce computation complexity.
PDF KSCI

Knowledge-Based Web Document Filtering (지식기반 웹 문서 필터링)

황상규;김상모;변영태
- Proceedings of the Korean Information Science Society Conference
- /
- 1999.10b
- /
- pp.51-53
- /
- 1999
인터넷에서 검색 가능한 정보의 양은 폭발적으로 증가하고 있으며, 그에 따라 웹 기반 정보검색시스템은 사용자가 원하는 정보만을 필터링하여 이용자의 정보검색 수행과정에 부담을 덜어줄 필요가 있다. 본 연구에서는 웹 정보검색에 익숙치 못한 초보 이용자들이 실제 웹 정보검색을 수행하는데 있어 발생할 수 있는 문제점을 살펴보고, 초보 이용자들의 보다 편리한 웹 정보검색을 도와줄 수 있도록 하기 위하여 WordNet을 활용한 지식베이스와 SDCC(Semantic Distance for Common Category)를 이용한 웹 문서 필터링 알고리즘을 개발하고 그 효율성을 확인하였다.
PDF

LG HomeNet Solution 적용 사례

Park, Hyun
- Korea Information Processing Society Review
- /
- v.11 no.3
- /
- pp.91-94
- /
- 2004
최근 정통부는 9년 동안 정체되어 있는 국민소득 1만불 시대에서 2만불 시대로의 돌파를 위해서 "IT 839전략"을 추진하고 있다. 그 주요 요지는 8대 신규 서비스, 3대 인프라, 그리고 9대 신성장 엔진을 통해 2012년에 2만불 시대를 달성하자는 것이다. "IT386전략"의 핵심 내용을 살펴보면 항목 하나 하나가 홈네트워크, 더 나아가 유비쿼터스 네트워크(Ubiquitous Network)의 구성 요소로 가득 차 있다. 홈네트워크는 PC, 인터넷, 모바일 이후를 대표하는 차세대 IT Key word로서 자리 매김하고 있으며 시장 규모적인 측면이나, 국민경제에 미치는 파급효과, 각 개인의 생활의 변화 등 다방면에서 큰 파장을 일으킬 것으로 예상되고 있다. (중략)으킬 것으로 예상되고 있다. (중략)
PDF

Web Document Clustering Using Statistical Techniques & Tag Information on the Specific-Domain Web site (전문 웹 사이트에서의 통계적 기법과 태그 정보를 이용한 문서 분류)

조은휘;변영태
- Proceedings of the Korea Inteligent Information System Society Conference
- /
- 2002.11a
- /
- pp.297-302
- /
- 2002
특정 영역에 대해 사용자에게 관련 정보를 제공하는 서비스를 위해 정보 에이전트를 개발하고 있다. 이 시스템은 웹 상에서 문서를 수집해 오는데 특정 영역과 관련한 지식베이스를 토대로 하고 있는데, 이들 중 몇몇 전문 사이트 내의 정보가 많이 포함되어 있음을 볼 수 있다. 그러므로 전문 사이트 내의 관련 문서 수집은 중요한 의의가 있다. 본 논문에서는 이들 전문 사이트 내의 전문 문서 수집을 위해 문서간의 유사성을 토대로 클러스터링 한다. 즉, 문서내의 텀(term)과 HTML 태그(tag), 지식베이스의 WordNet 계층구조를 data로 하고 SVD(Singular Value Decomposition)을 사용하여 문서간의 관계를 밝혀내었다.
PDF

Korean Isolated Word Recognition Using Modular Structured Neural Network (모듈구조 신경망을 이용한 한국어 단어 인식에 관한 연구)

최환진
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1991.06a
- /
- pp.11-14
- /
- 1991
음소단위로 구성된 음소군들 각각에 대해 구성된 신경 회로망을 하나로 통합하는 모듈구조로 신경망을 이용하여 일반적인 예약 시스템에서 사용할 수 있는 어휘인 시간명, 월명, 지역명등 총 34 단어에 대한 인식 실험내용을 기술한다. 구문회로망(context net)를 이용하는 경우에 약 91.2%의 인식율을, 단순히 음소단위를 기반으로하여 인식할 경우에 약 72%의 인식율을 얻으므로써, 음소 단위 인식시스템의 경우에 보다 높은 인식율을 얻기 위해서는 상위 level의 처리가 수반되어야 함을 확인할 수 있었다.
PDF

Improvement of a Sentence Analysis System through Lexical Expansion (어휘확장을 통한 문장분석 시스템의 개선)

Kim Min-Chan;Kim Gon;Bae Jae-Hak
- Proceedings of the Korean Information Science Society Conference
- /
- 2005.07b
- /
- pp.496-498
- /
- 2005
본 논문에서는 미등록 어휘로 인한 구문분석의 실패를 해결하는 방법으로 WordNet의 유의어 정보를 이용하였다. 이 방법을 또한 설화용 온톨러지 OfN의 어휘확장에 적용하였다. 실험을 통하여 구문분석 과정에서 나타나는 미등록 어휘문제의 해결과 문장의 의미분석 과정이 순조롭게 진행될 수 있음을 확인하였다.
PDF

Unsupervised Noun Sense Disambiguation using Local Context and Co-occurrence (국소 문맥과 공기 정보를 이용한 비교사 학습 방식의 명사 의미 중의성 해소)

Lee, Seung-Woo;Lee, Geun-Bae
- Journal of KIISE:Software and Applications
- /
- v.27 no.7
- /
- pp.769-783
- /
- 2000
In this paper, in order to disambiguate Korean noun word sense, we define a local context and explain how to extract it from a raw corpus. Following the intuition that two different nouns are likely to have similar meanings if they occur in the same local context, we use, as a clue, the word that occurs in the same local context where the target noun occurs. This method increases the usability of extracted knowledge and makes it possible to disambiguate the sense of infrequent words. And we can overcome the data sparseness problem by extending the verbs in a local context. The sense of a target noun is decided by the maximum similarity to the clues learned previously. The similarity between two words is computed by their concept distance in the sense hierarchy borrowed from WordNet. By reducing the multiplicity of clues gradually in the process of computing maximum similarity, we can speed up for next time calculation. When a target noun has more than two local contexts, we assign a weight according to the type of each local context to implement the differences according to the strength of semantic restriction of local contexts. As another knowledge source, we get a co-occurrence information from dictionary definitions and example sentences about the target noun. This is used to support local contexts and helps to select the most appropriate sense of the target noun. Through experiments using the proposed method, we discovered that the applicability of local contexts is very high and the co-occurrence information can supplement the local context for the precision. In spite of the high multiplicity of the target nouns used in our experiments, we can achieve higher performance (89.8%) than the supervised methods which use a sense-tagged corpus.
PDF

Search Result 258, Processing Time 0.035 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)