Search | Korea Science

Experiments on Pseudo Relevance Feedback in Probabilistic Information Retrieval Model (확률적 정보 검색 모델에서의 유사 적합성 피드백 실험)

Cho, Bong-Hyun;Lee, Chang-Kee;An, Joo-Hui;Lee, Gary Geun-Bae
- Annual Conference on Human and Language Technology
- /
- 2001.10d
- /
- pp.183-190
- /
- 2001
본 논문은 확률기반 자연어 검색 시스템 POSNIR/E를 이용한 여러 가지 유사 적합성 피드백 방법들이 검색 시스템의 성능 향상에 기여할 수 있는 정도를 보여주고, 확률 기반 정보 검색 시스템에 적합한 유사 적합성 피드백 수행 방법을 제시한다. POSNIR/E는 한국어 자연어 검색 시스템, POSNIR를 기반으로 만들어진 영어 자연어 검색 시스템이다. 이 시스템은 성능 향상을 위한 질의 확장의 방법으로 검색 단계에서 유사 적합성 피드백을 사용한다. 검색 단계에서 영어 태거에 의해 태깅된 사용자 질의로부터 질의어를 추출하고 초기 검색을 수행한다. 유사 적합성 피드백을 위하여 초기 검색 결과 중 상위 5개의 문서에 나타나는 키워드를 중요도에 따라 내림차순 정렬하여 상위 10개의 키워드를 초기 질의어에 확장한다. 이렇게 확장된 질의어로 최종 검색을 수행한다. TREC 평가용 테스트 컬렉션 WT10g와 TREC-9의 질의 적합문서 집합을 이용하여 여러 가지 TSV 함수를 사용하여 검색 성능을 평가 하였다. 실험 결과 유사 적합성 피드백을 사용할 경우 TSV 함수에 확률 모델의 CF 요소 뿐만 아니라 TF 요소 등을 적용 시킬 경우 성능 향상에 기여할 수 있음을 알 수 있었다. 또한 색인어와 검색어로 단일어 뿐만 아니라 복합어도 사용할 경우 성능이 향상됨을 알 수 있다.
PDF

Building Domain Ontology through Concept and Relation Classification (개념 및 관계 분류를 통한 분야 온톨로지 구축)

Huang, Jin-Xia;Shin, Ji-Ae;Choi, Key-Sun
- Journal of KIISE:Software and Applications
- /
- v.35 no.9
- /
- pp.562-571
- /
- 2008
For the purpose of building domain ontology, this paper proposes a methodology for building core ontology first, and then enriching the core ontology with the concepts and relations in the domain thesaurus. First, the top-level concept taxonomy of the core ontology is built using domain dictionary and general domain thesaurus. Then, the concepts of the domain thesaurus are classified into top-level concepts in the core ontology, and relations between broader terms (BT) - narrower terms (NT) and related terms (RT) are classified into semantic relations defined for the core ontology. To classify concepts, a two-step approach is adopted, in which a frequency-based approach is complemented with a similarity-based approach. To classify relations, two techniques are applied: (i) for the case of insufficient training data, a rule-based module is for identifying isa relation out of non-isa ones; a pattern-based approach is for classifying non-taxonomic semantic relations from non-isa. (ii) For the case of sufficient training data, a maximum-entropy model is adopted in the feature-based classification, where k-NN approach is for noisy filtering of training data. A series of experiments show that performances of the proposed systems are quite promising and comparable to judgments by human experts.
PDF KSCI

Semantic Information Retrieval using User-Word Intelligent Network (사용자 어휘지능망을 이용한 의미적 정보검색)

Kim, Chang-Hwan;Im, Ji-Hui;Choe, Ho-Seop;Yoon, Hwa-Mook;Ock, Cheol-Young
- Proceedings of the Korea Information Processing Society Conference
- /
- 2006.11a
- /
- pp.157-160
- /
- 2006
웹 자원이 방대함에 따라, 사용자가 원하는 정보를 얼마나 정확하게 제시하느냐가 정보검색시스템 성능을 판단하는 기준이 된다. 그러나 동형이의어만을 질의어로 이용한 검색 결과는 동형이의어 각 의미에 관련된 문서가 혼재되어 있거나, 특정 의미에 관련된 문서가 집중적으로 나타나는 현상을 볼 수 있다. 이에 본 논문에서는 한국어 사용자 어휘지능망(U-WIN)의 관계정보를 이용하여 질의어의 모호성을 해결하고 의미적 정보검색의 기반을 마련하고자 한다. 우선, 전문분야에 주로 사용되는 동형이의어와 보편적으로 사용하는 동형의어를 구번하여 질의어로 선정하고, '질의어+상위어' 형태의 확장 질의어에 대해 두 개의 포탈사이트(Google, Naver)를 대상으로 웹 문서를 검색하여 정확률이 각각 81.5%(Naver), 65.5%(Google)로 나타났다.
PDF

Ontology Mapping using Semantic Relationship Set of the WordNet (워드넷의 의미 관계 집합을 이용한 온톨로지 매핑)

Kwak, Jung-Ae;Yong, Hwan-Seung
- Journal of KIISE:Databases
- /
- v.36 no.6
- /
- pp.466-475
- /
- 2009
Considerable research in the field of ontology mapping has been done when information sharing and reuse becomes necessary by a variety of ontology development. Ontology mapping method consists of the lexical, structural, instance, and logical inference similarity computing. Lexical similarity computing used in most ontology mapping methods performs an ontology mapping by using the synonym set defined in the WordNet. In this paper, we define the Super Word Set including the hypenym, hyponym, holonym, and meronym set and propose an ontology mapping method using the Super Word Set. The results of experiments show that our method improves the performance by up to 12%, compared with previous ontology mapping method.
PDF KSCI

A Study on A Korean Noun Semantic TAG based on Semantic Features (의미속성에 기반한 한국어 명사 의미 TAG에 관한 연구)

Lee, S.;Cho, P.;Ahn, M.;Ock, C.;Park, J.;Park, D.
- Annual Conference on Human and Language Technology
- /
- 1998.10c
- /
- pp.412-418
- /
- 1998
의미 TAG는 한국어 기초어휘에 대한 개념지식을 구축하는 데 기본이 될 뿐만 아니라, 문장 분석시의 구조적 모호성과 단어 의미 모호성을 해소하는 중요한 단서를 제공할 수 있다. 이러한 의미 TAG가 실용적으로 여러 응용 시스템에서 사용되기 위해서는 광범위하고 타당한 자료를 바탕으로 하여 객관적인 방법으로 설정 되어야 한다. 국어사전의 뜻풀이말에서의 상위개념을 표제어의 상위어로 선정하는 bottom-up 방식으로 구축하였던 한국어 명사의미체계는 근본적으로 사전편찬자의 비일관적인 뜻풀이말의 기술에 따른 여러 문제점이 있었다. 본 연구에서는 이러한 문제점들을 해결하기 위해서 사전 뜻풀이말에서 상위개념을 수식하는 어절과 용언의 의미호응관계에서 상위개념의 의미속성을 추출하고, 이들 의미속성에 의한 명사 의미체계를 구축하여 이를 바탕으로 명사의미 TAG를 설정할 수 있도록 하였다.
PDF

Analysis of Keywords and Language Networks of Pedagogical Problems in the Secondary-School Teacher's Employment Exam : Focusing on the 2019~2022 School Year Exam

Kwon, Choong-Hoon
- Journal of the Korea Society of Computer and Information
- /
- v.27 no.7
- /
- pp.115-124
- /
- 2022
The purpose of this study is to analyze and present keywords, trends, and language networks of keywords for each year of the pedagogical exam of the secondary teacher's employment exam for the 2019~2022 school year. The main research methods were text mining technique and language network analysis method, and analysis programs were KrKwic, Wordcloud Maker, Ucinet6, NetDraw, etc. The research results are as follows; First, keywords such as teacher, student, curriculum, class, and evaluation appeared in the top rankings, and keywords (online, wiki, discussion ceremony, information, etc.) that reflect the recent online class progress in the current COVID-19 situation also tended to appear. The keywords with high frequency of occurrence in the four-year integrated text were student(44), teacher(39), class(27), school(18), curriculum(16), online(10), and discussion method(8). Second, the overall language network of the keywords with high frequency of 4 years showed a significant level of density(0.566), total number of links(492), and average degree of links(16.4). The degree centrality was found in the order of teacher(199.0), class(197.0), student(185.0), and school(150.0). Betweenness centrality was found in the order of teacher(30.859), class(18.956), student(16.054), and school (15.745). It is expected that the results of this study will serve as data to be considered for preparatory teachers, institutions and related persons, and teachers and administrators of secondary school teacher training institutions.
https://doi.org/10.9708/jksci.2022.27.07.115 인용 PDF KSCI HTML

The Method of Deriving Keywords Using Concept Rules (개념 규칙을 이용한 키워드 도출방법)

이태헌;박기홍
- Proceedings of the Korean Information Science Society Conference
- /
- 2002.10d
- /
- pp.685-687
- /
- 2002
일반적으로 인간이 사용하는 몇 개의 주요단어를 이용하여, 문서의 분야나 주제어가 되는 일본어 키워드를 추출하는 점에 주목한다. 먼저, 학술논문에서 저자 자신이 부여한 키워드 중 분야 명이나 주제어가 문서 중에 출현하지 않는 경우를 분석하고, 단어의 개념정보를 기초로 복합어 생성규칙을 구축한다. 문서 의미와 상관없는 키워드의 추출을 억제하기 위해 중요도 결정법을 새롭게 제안한다. 추출된 키워드의 타당성 검사를 위해 자연.음성언어에 관한 일본어 논문 65파일의 타이틀과 초록부분을 이용하여 추출된 키워드의 타당성에 대한 실험을 한 결과 추출 정밀도는 중요도의 상위 1개를 출력한 경우 75%가 되어 제안방법의 유효성을 확인할 수 있었다.
PDF

Dictionary Making for Disambiguation (동사의 애매성 해소를 위한 구문의미사전의 구축)

Song, Young-Bin;Chae, Young-Soog;Park, Yong-Il;Lee, Jun-Min;Seol, Kah-Young;Hwang, Hye-Ri;Han, Na-Ri;Choi, Key-Sun
- Annual Conference on Human and Language Technology
- /
- 1999.10e
- /
- pp.280-287
- /
- 1999
동사의 애매성이란 동일 동사 내부에서 공기하는 명사의 상충적 의미의 분포에 의해 발생한다. 이는 동일한 동사라 하더라도 명사의 상위개념, 흑은 개개의 명사에 따라 동사의 의미가 달라진다는 것을 의미한다. 동사의 애매성 해소를 위한 구문의미사전은 동사가 갖는 격틀과 논항에 오는 명사의 단어 집합에 의해 구성된다. 기계용 사전에서의 동사의 애매성이란 명사의 상위개념, 혹은 개개의 명사에 관한 정보가 결여될 때 나타난다. 지금까지의 구문의미사전은 개개의 동사가 갖는 격틀을 중심으로 논합명사의 예만을 제시하거나 명사의 상위개념을 기술하는 형식으로 구성되어 왔다. 이는 형식적인 패턴의 추출에는 유용하지만 대역어 선정을 위한 구문의미사전과 같은 섬세한 의미 정보를 필요로 하는 사전에서는 거의 효력을 발휘하지를 못한다. 다국어를 전제로 한 동사 대역어의 추출을 목적으로 하는 구문의미사전에서는 동사와 공기하는 논항명사의 철저한 추출과 검증에 의한 명사목록의 구축이 애매성 해소와 정확한 동사 대역어의 선정에 전제가 된다. 본 논문에서는 KAIST Corpus를 기반으로 현재 구축 중인 한국어 구문의미사전의 개요와 구축 과정에서 얻어진 방법론을 소개한다. 이 연구개발 결과는 과학기술부 KISTEP 특정연구개발과제 핵심소프트웨어개발 국어정보처리기술개발 중 "대용량 국어정보 심층 처리 및 품질 관리 기술 개발"의 지원을 받았다.
PDF

Detection of Character Emotional Type Based on Classification of Emotional Words at Story (스토리기반 저작물에서 감정어 분류에 기반한 등장인물의 감정 성향 판단)

Baek, Yeong Tae
- Journal of the Korea Society of Computer and Information
- /
- v.18 no.9
- /
- pp.131-138
- /
- 2013
In this paper, I propose and evaluate the method that classifies emotional type of characters with their emotional words. Emotional types are classified as three types such as positive, negative and neutral. They are selected by classification of emotional words that characters speak. I propose the method to extract emotional words based on WordNet, and to represent as emotional vector. WordNet is thesaurus of network structure connected by hypernym, hyponym, synonym, antonym, and so on. Emotion word is extracted by calculating its emotional distance to each emotional category. The number of emotional category is 30. Therefore, emotional vector has 30 levels. When all emotional vectors of some character are accumulated, her/his emotion of a movie can be represented as a emotional vector. Also, thirty emotional categories can be classified as three elements of positive, negative, and neutral. As a result, emotion of some character can be represented by values of three elements. The proposed method was evaluated for 12 characters of four movies. Result of evaluation showed the accuracy of 75%.
https://doi.org/10.9708/jksci.2013.18.9.131 인용 PDF KSCI

A Framework for WordNet-based Word Sense Disambiguation (워드넷 기반의 단어 중의성 해소 프레임워크)

Ren, Chulan;Cho, Sehyeong
- Journal of the Korean Institute of Intelligent Systems
- /
- v.23 no.4
- /
- pp.325-331
- /
- 2013
This paper a framework and method for resolving word sense disambiguation and present the results. In this work, WordNet is used for two different purposes: one as a dictionary and the other as an ontology, containing the hierarchical structure, representing hypernym-hyponym relations. The advantage of this approach is twofold. First, it provides a very simple method that is easily implemented. Second, we do not suffer from the lack of large corpus data which would have been necessary in a statistical method. In the future this can be extended to incorporate other relations, such as synonyms, meronyms, and antonyms.
https://doi.org/10.5391/JKIIS.2013.23.4.325 인용 PDF KSCI

Search Result 161, Processing Time 0.045 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)