Search | Korea Science

Automatic Text Categorization by Term Weighting and Inverted Category Frequency (용어 가중치와 역범주 빈도에 의한 자동문서 범주화)

Lee, Kyung-Chan;Kang, Seung-Shik
- Annual Conference on Human and Language Technology
- /
- 2003.10d
- /
- pp.14-17
- /
- 2003
문서의 확률을 이용하여 자동으로 문서를 분류하는 문서 범주화 기법의 대표적인 방법이 나이브 베이지언 확률 모델이다. 이 방법의 기본 형식은 출현 용어의 확률 계산 방법이다. 하지만 실제 문서 범주화 과정에서 출현하지 않는 용어들도 성능에 많은 영향을 줄 수 있으며, 출현 용어들에 대한 빈도 이외의 역범주 빈도나 용어가중치를 적용하여 문서 범주화 시스템의 성능을 향상시킬 수 있다. 본 논문에서는 나이브 베이지언 확률 모델에 출현 용어와 출현하지 않는 용어들에 대한 smoothing 기법을 적용하여 실험하였다. 성능 평가를 위해 뉴스그룹 문서들을 이용하였으며, 역범주 빈도와 가중치를 적용했을 때 나이브 베이지언 확률 모델에 비해 약 7% 정도 성능 개선 효과가 있었다.
PDF

The Term and Classification of Structure System with Non-rigid Member (연성구조시스템의 분류체계와 용어)

Lee, Ju-Na;Park, Sun-Woo;Kim, Seung-Deog;Park, Chan-Soo
- Journal of Korean Association for Spatial Structures
- /
- v.4 no.2 s.12
- /
- pp.99-105
- /
- 2004
The structure systems with non-rigid member were classified by the composition type of line and surface members. As a result of the classification, there are 1-way cable structure, cable net and radial cable net structure in the line member system. And there are pneumatic structure and suspension membrane structure in surface member system. In addition, when the line and surface members are composed together, there is the hybrid membrane system which are divided into hanging type and supported type. In this paper, the Korean terms of structure systems with non-rigid member are recommended through this classification.
PDF

A Junkmail Checking System Using Fuzzy Relational Products (퍼지 관계 곱을 이용한 정크메일 분류 시스템)

박정선;김창민;김용기
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 2001.12a
- /
- pp.341-344
- /
- 2001
20세기 후반 인터넷의 발전을 기반으로 전자메일은 현재의 대표적인 개인간 정보전달 수단으로 자리 잡게 되었다. 그러나 전자메일 사용자들은 인터넷상에 개인 전자메일 주소가 노출되므로 해서 많은 정크메일(junkmail)을 수신하게 되었는데, 정크메일이란 기업의 광고 선전물과 같이 수신을 원하지 않는 전자메일을 의미한다. 이러한 정크메일의 증가에 따라 정크메일을 분류하는 수단이 필요하게 되었는데, 현재까지는 사용자가 입력한 송신자의 전자메일 주소 또는 도메인 주소를 등록하여 차단하거나 제목에 특정 단어를 포함한 메일을 완전히 삭제하여 버리는 기술수준에 머무르고 있다. 본 논문에서는 퍼지 관계 곱을 기반으로 메일의 내용에 의미적으로 접근하여 정크메일을 분류하는 시스템을 제안한다. 이는 퍼지 관계곱 연산을 이용하여 미리 정의한 정크용어들과 사용자에게 수신되는 전자메일 내의 용어들간 의미적 포함관계를 분석하고 그를 통해 전자메일의 정크도(degree of junk)를 추출한다. 각 전자메일별로 추출된 정크도는 사용자가 부여하는 정크 기준치(SVJ, Standard Value of Junk)를 기분으로 정크메일과 비 정크메일로 분류한다. 제안된 기법은 사용자가 특정 개수의 동일한 전자메일에 대해 느끼는 정크도를 기준으로 분류한 정크메일 수를 비교하여 그 효용성을 증명하였다.
PDF

A HS tariff classification service based on a knowledge convergence performance system supporting decision elements and field terms (결정요소 및 현장용어 지원 지식융합 수행 시스템 기반의 HS 관세분류 서비스에 관한 실증 연구)

Kim, Eunsoo;Song, ByungJun;Lee, Jong Yun
- Journal of the Korea Convergence Society
- /
- v.6 no.1
- /
- pp.49-55
- /
- 2015
In the FTA environment, it is necessary to comply with the rules of origin in order to receive duty-free benefits. To do this, they have to precede the Harmonized System(HS) tariff classification of the goods and understand thoroughly the basic principles that constitute the tariff schedule of HS classification. For the correct classification, they should understand exactly the product name of "Heading" about the items, "Legal Note" in the relevant "Section" or "Chapter" as well as provisions of the commentary. Therefore, this paper proposes to develop a HS classification services based on the performance system of knowledge convergence of field terms commonly used in various industries. In result, our services can provide users the conveniences which users first selects one of seven decision elements of the classification and perform the classification easily and accurately.
https://doi.org/10.15207/JKCS.2015.6.1.049 인용 PDF KSCI

An Experimental Study on Opinion Classification Using Supervised Latent Semantic Indexing(LSI) (지도적 잠재의미색인(LSI)기법을 이용한 의견 문서 자동 분류에 관한 실험적 연구)

Lee, Ji-Hye;Chung, Young-Mee
- Journal of the Korean Society for information Management
- /
- v.26 no.3
- /
- pp.451-462
- /
- 2009
The aim of this study is to apply latent semantic indexing(LSI) techniques for efficient automatic classification of opinionated documents. For the experiments, we collected 1,000 opinionated documents such as reviews and news, with 500 among them labelled as positive documents and the remaining 500 as negative. In this study, sets of content words and sentiment words were extracted using a POS tagger in order to identify the optimal feature set in opinion classification. Findings addressed that it was more effective to employ LSI techniques than using a term indexing method in sentiment classification. The best performance was achieved by a supervised LSI technique.
https://doi.org/10.3743/KOSIM.2009.26.3.451 인용 PDF

International Patent Classificaton Using Latent Semantic Indexing (잠재 의미 색인 기법을 이용한 국제 특허 분류)

Jin, Hoon-Tae
- Proceedings of the Korea Information Processing Society Conference
- /
- 2013.11a
- /
- pp.1294-1297
- /
- 2013
본 논문은 기계학습을 통하여 특허문서를 국제 특허 분류(IPC) 기준에 따라 자동으로 분류하는 시스템에 관한 연구로 잠재 의미 색인 기법을 이용하여 분류의 성능을 높일 수 있는 방법을 제안하기 위한 연구이다. 종래 특허문서에 관한 IPC 자동 분류에 관한 연구가 단어 매칭 방식의 색인 기법에 의존해서 이루어진바가 있으나, 현대 기술용어의 발생 속도와 다양성 등을 고려할 때 특허문서들 간의 관련성을 분석하는데 있어서는 단어 자체의 빈도 보다는 용어의 개념에 의한 접근이 보다 효과적일 것이라 판단하여 잠재 의미 색인(LSI) 기법에 의한 분류에 관한 연구를 하게 된 것이다. 실험은 단어 매칭 방식의 색인 기법의 대표적인 자질선택 방법인 정보획득량(IG)과 카이제곱 통계량(CHI)을 이용했을 때의 성능과 잠재 의미 색인 방법을 이용했을 때의 성능을 SVM, kNN 및 Naive Bayes 분류기를 사용하여 분석하고, 그중 가장 성능이 우수하게 나오는 SVM을 사용하여 잠재 의미 색인에서 명사가 해당 용어의 개념적 의미 구조를 구축하는데 기여하는 정도가 어느 정도인지 평가함과 아울러, LSI 기법 이용시 최적의 성능을 나타내는 특이값의 범위를 실험을 통해 비교 분석 하였다. 분석결과 LSI 기법이 단어 매칭 기법(IG, CHI)에 비해 우수한 성능을 보였으며, SVM, Naive Bayes 분류기는 단어 매칭 기법에서는 비슷한 수준을 보였으나, LSI 기법에서는 SVM의 성능이 월등이 우수한 것으로 나왔다. 또한, SVM은 LSI 기법에서 약 3%의 성능 향상을 보였지만 Naive Bayes는 오히려 20%의 성능 저하를 보였다. LSI 기법에서 명사가 잠재적 의미 구조에 미치는 영향은 모든 단어들을 내용어로 한 경우 보다 약 10% 더 향상된 결과를 보여주었고, 특이값의 범위에 따른 성능 분석에 있어서는 30% 수준에 Rank 되는 범위에서 가장 높은 성능의 결과가 나왔다.
https://doi.org/10.3745/PKIPS.y2013m11a.1294 인용 PDF

A Hypertext Categorization Model Exploiting Link and Incrementally Available Category Information (점진적으로 계산되는 분류정보와 링크정보를 이용한 하이퍼텍스트 문서 분류 모델)

Oh, Hyo-Jung;Lim, Jeong-Mook;Lee, Mann-Ho;Myaeng, Sung-Hyon
- Annual Conference on Human and Language Technology
- /
- 1999.10e
- /
- pp.89-96
- /
- 1999
본 논문은 하이퍼텍스트가 갖는 중요한 특성인 링크 정보를 활용한 문서 분류 모델을 제안한다. 하이퍼링크는 문서간의 관계를 나타내는 유용한 정보로서 링크를 통해 연결된 두 문서는 내용적으로 관련이 있어 검색에 도움을 준다는 것은 이미 밝혀진바 있다. 본 논문에서는 이러한 과거 연구를 바탕으로 새로운 문서 분류 모델을 제안하는데, 이 모델의 주안점은 대상 문서와 링크로 연결된 이웃 문서의 내용 및 범주를 분석하여 대상 문서 벡터를 조정하고, 이를 근거로 문서의 범주를 결정한다. 이웃 문서에 포함된 용어를 반영함으로써 대상 문서의 내용을 확장 해석하고, 이웃 문서의 가용 분류 정보가 있는 경우 이를 참조함으로써 정확도 향상을 기한다. 이 모델은 이웃한 문서의 범주가 미리 할당되어 있지 않은 경우 용어 기반 분류 방법으로 가용 범주를 할당하고, 이렇게 할당된 분류 정보가 다시 새로운 문서의 범주를 결정할 때 사용됨으로써, 문서 집합 전체의 분류가 점진적으로 이루어지며 그 정확도를 더해 나가는 효과를 가져올 수 있다. 이러한 접근 방법은 일반 웹 환경에 적용할 수 있는데, 특히 하이퍼텍스트를 주제별로 분류하여 관리하는 검색 엔진의 경우 매일 쏟아져 나오는 새로운 문서와 기존 문서간의 링크를 활용함으로써 전체 시스템의 점진적인 분류에 매우 유용하다. 제안된 모델을 검증하기 위하여 Reuter-21578과 계몽사(ETRI-Kyemong) 자료를 대상으로 실험한 결과 18.5%의 성능 향상을 얻었다.
PDF

Domain-specific Ontology Construction by Terminology Processing (전문용어의 처리에 의한 도메인 온톨로지의 구축)

임수연;송무희;이상조
- Journal of KIISE:Software and Applications
- /
- v.31 no.3
- /
- pp.353-360
- /
- 2004
Ontology defines the terms used in a specific domain and the relationships between them and represents them as hierarchical taxonomy. The present paper proposes a semi-automatic domain-specific ontology construction method based on terminology Processing. For this purpose, it presents an algorithm to extract terminology according to the noun/suffix pattern of terminology in domain texts and find their hierarchical structure. The experiment was carried out using pharmacy-related documents. As singleton terminology with noun/suffix were identified, the average accuracy was 92.57%. In case of multi-word terminology, the average accuracy was 66.64%. The constructed ontology forms natural semantic clusters with based on suffices and semantic information, so can be utilized in approaches to specific knowledge such as information look-up or as the base of inference to improve searching abilities.
PDF KSCI

Performance Evaluation for Word Clustering (용어 클러스터링의 성능 평가)

Park, Eun-Jin;Kim, Jae-Hoon;Ock, Cheol-Young
- Annual Conference on Human and Language Technology
- /
- 2005.10a
- /
- pp.43-49
- /
- 2005
이 논문에서는 전자 사전의 뜻 풀이말을 이용하여 용어를 자동 분류하는 용어 클러스터링 시스템을 설계하였다. 클러스터링 성능에 영향을 미치는 요소로 자질 선택 자질 표현 그리고 유사도 측정 등이 있다. 이 논문에서는 이러한 요소들이 용어 클러스터링에 미치는 영향을 평가해보았다. 클러스터링 결과를 객관적으로 비교하기 위해서 용어 클러스터링 결과와 한국어 의미 계층망에서 추출한 정답 클러스터를 비교하였다 실험 결과, 용어의 뜻 풀이말만 자질로 사용한 방법보다는 뜻 풀이말 자질을 확장하는 방법이 훨씬 더 좋은 결과를 보였다.
PDF

Design and Implementation of Extracting Medical Terms Using Web Mining (웹 마이닝을 활용한 의학 용어 추출 시스템 설계 및 구현)

Choi, Wook-Hwan;Shin, Jung-Hoon;Lee, Sang-Jun
- Proceedings of the Korean Information Science Society Conference
- /
- 2011.06c
- /
- pp.56-59
- /
- 2011
최근 방대해진 의료 정보 관리와 통합을 위해 다양한 전자기록 시스템이 개발되어왔다. 이 중에 EMR(Electronic Medical Record)은 병원 내의 의료 정보를 전산으로 처리하는 것이다. 이때 많은 의학 용어가 사용되는데 이것을 체계적으로 지원하기 위해 SNOMED(Systematized Nomenclature of Medicine)와 같은 용어 체계가 필요하다. 본 논문에서는 이러한 용어 체계를 위해 의학 용어를 자동으로 구축하는 시스템을 제안하고자 한다. 제안한 시스템은 웹 마이닝을 통해 자동으로 웹에서 데이터를 수집하고 분류해서 의학 용어 데이터를 구축하게 된다.

Search Result 482, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)