Search | Korea Science

Pre-Processing of Korean Syntactic Analyzer for Korean to English MT (한영 자동 번역을 위한 한국어 구문 분석 전처리)

김영길;양성일;서영애;김창현;홍문표;최승권
- Proceedings of the Korean Information Science Society Conference
- /
- 2001.10b
- /
- pp.175-177
- /
- 2001
형태소 해석 결과 생성되는 형태소 옅은 구문 분석을 수행하기에는 적절하지 않은 구문 단위로 구성되어 있는 경우가 많으며 이로 인해 구문 분석기가 불필요한 연산을 수행하여 과도한 구문 트리를 생성하는 원인이 된다. 따라서 본 논문에서는 한영 자동 번역의 한국어 구문 분석기 성능 향상 및 자연스러운 대역문 생성을 위하여 시간 부사구와 명사구에 대한 구묶음을 위한 구문 분석 전처리 방법을 제안하며 이를 위한 각 구 단위의 대역 패턴을 정의한다. 방송자막 및 매뉴얼 문장을 대상으로 실험한 결과, 각 문장 구문 단위를 평균적으로 26% 정도 감소시킴으로써 불필요한 파스 트리의 생성을 배제하여 구문 분석기의 성능을 향상시킬 수 있었다.
PDF

An n-gram-based Indexing Method for Effective Retrieval of Hangul Texts (한글 문서의 효과적인 검색을 위한 n-gram 기반의 색인 방법)

이준호;안정수;박현주;김명호
- Journal of the Korean Society for information Management
- /
- v.13 no.1
- /
- pp.47-63
- /
- 1996
Conventional automatic indexing methods for Hangul texts can be classified into two groups as follows: One is to extract index terms by removing non-indexable segments from word-phrases, and the other is to generate index terms from the morphemes of word-phrases. The former suffers from the problem of word boundaries when documents contain many compound nouns. The latter can overcome the word boundary problem by extracting simple nouns, but has many overheads to develop a lot of linguistic knowledges needed in the indexing procedure. In this paper we propose a new indexing method based on n-grams. This method alleviates the problems of previous indexing methods related with word boundaries and linguistic knowledges. We also compare the effectiveness of the n-gram based indexing method with that of the previous ones.
PDF

Splitting Algorithms and Recovery Rules for Zero Anaphora Resolution in Korean Complex Sentences (한국어 복합문에서의 제로 대용어 처리를 위한 분해 알고리즘과 복원규칙)

Kim, Mi-Jin;Park, Mi-Sung;Koo, Sang-Ok;Kang, Bo-Yeong;Lee, Sang-Jo
- Journal of KIISE:Software and Applications
- /
- v.29 no.10
- /
- pp.736-746
- /
- 2002
Zero anaphora occurs frequently in Korean complex sentences, and it makes the interpretation of sentences difficult. This paper proposes splitting algorithms and zero anaphora recovery rules for the purpose of handling zero anaphora, and also presents a resolution methodology. The paper covers quotations, conjunctive sentences and embedded sentences out of the complex sentences shown in the newspaper articles, with an exclusion of embedded sentences of auxiliary verb. We manage the quotations using the equivalent noun phrase deletion rule according to subject person constraint, the nominalized embedded sentences using the equivalent noun phrase deletion rule, the adnominal embedded sentences using the relative noun phrase deletion rule and the conjunctive sentences using the conjunction reduction rule in reverse. The classified table of the endings which relate to a formation of the complex sentences is used for splitting the complex sentences, and the syntactic rules, applied when being omitted, are used in reverse for recovering zero anaphora. The presented rule showed the result of 83.53% in perfect resolution and 11.52% in partial resolution.
PDF KSCI

Word Sense Disambiguation Based on Local Syntactic Relations and Sense Co-occurrence Information (국소 구문 관계 및 의미 공기 정보에 기반한 명사 의미 모호성 해소)

Kim, Young-Kil;Hong, Mun-Pyo;Kim, Chang-Hyun;Seo, Young-Ae;Yang, Seong-Il;Ryu, Chul;Huang, Yin-Xia;Choi, Sung-Kwon;Park, Sang-Kyu
- Annual Conference on Human and Language Technology
- /
- 2002.10e
- /
- pp.184-188
- /
- 2002
본 논문에서는 단순히 주변에 위치하는 어휘들간의 문맥 공기 정보를 이용하는 방식과는 달리 국소 구문 관계 및 의미 공기 정보에 기반한 명사 의미 모호성 해소 방안을 제안한다. 기존의 WSD 방법은 구조 분석의 어려움으로 인하여 문장의 구문 관계를 충분히 고려하지 못하고 주변 어휘들과의 공기 관계로 그 의미를 파악하려 했다. 그러나 본 논문에서는 동사구의 논항 의미 관계뿐만 아니라 명사구내에서의 의미 관계도 고려한 국소 구문관계를 고려한 명사 의미 모호성 해소 방법을 제안한다. 이 때, 명사들의 의미는 자동번역 시스템의 목적에 맞게 공기(co-occurrence)하는 동사들에 따라 분류하였다. 그리고 한중 자동 번역 지식으로 사용되는 명사 의미 코드가 부착된 74,880 의미 격틀의 의미 공기정보를 이용하였으며 형태소 태깅된 말뭉치로부터 의미모호성이 발생하지 않게 의미 공기정보 및 명사구 의미 공기 정보를 자동으로 추출하였다. 실험 결과, 의미 모호성이 발생하는 명사들에 대해서 83.9%의 의미 모호성 해소 정확률을 보였다.
PDF

Syntactic and Semantic Analysis of Korean Verb 'Kat-' ('같다' 구문의 통사.의미적 특성)

Nam, Yun-Jin;Han, Young-Gyun
- Annual Conference on Human and Language Technology
- /
- 1992.10a
- /
- pp.385-402
- /
- 1992
용언 '같다'는 다양한 의미를 지니는데, 그 가운데 [동일]이나 [유사]를 나타내는 '같다' 구문은 '비교'의 논리가 적용되는 문장들로서 문장을 이루는 명사구의 의미 특성, 명사구 사이의 의미관계, 문장 유형등의 요소에 따라 의미 해석이 달라진다. 이 유형의 '같다' 구문은 특정 문형의 실현이 명사구들의 의미 관계에 따라 제약을 받으며, 또 실현되는 경우에도 [동일]이나 [유사]라는 [비교]의 의미를 갖지 못하고 [비유]의 의미를 나타내게 된다. 이러한 의미범주의 변화는, 특정조건하에서의 '비교'가 현실논리에서는 성립할 수 없는 반면 언어논리에서는 수용될 때 나타나는 두 논리간의 괴리를 보완하는 기제인 것으로 생각된다. 한편, [동일]이나 [유사]를 나타내는 '같다'와 [추측] 혹은 [불확실한 단정]을 나타내는 '같다'는 통사구조와 의미해석 논리에서 다른 양상을 보인다. 이들은 항상 '(-ㄴ/ㄹ) 것 같다'와 같은 구성양식을 갖는데, 그럼에도 불구하고 단문구조로 해석되는 것이다.
PDF

Mention Detection using Bidirectional LSTM-CRF Model (Bidirectional LSTM-CRF 모델을 이용한 멘션탐지)

Park, Cheoneum;Lee, Changki
- Annual Conference on Human and Language Technology
- /
- 2015.10a
- /
- pp.224-227
- /
- 2015
상호참조해결은 특정 개체에 대해 다르게 표현한 단어들을 서로 연관지어 주며, 이러한 개체에 대해 표현한 단어들을 멘션(mention)이라 하며, 이런 멘션을 찾아내는 것을 멘션탐지(mention detection)라 한다. 멘션은 명사나 명사구를 기반으로 정의되며, 명사구의 경우에는 수식어를 포함하기 때문에 멘션탐지를 순차 데이터 문제(sequence labeling problem)로 정의할 수 있다. 순차 데이터 문제에는 Recurrent Neural Network(RNN) 종류의 모델을 적용할 수 있으며, 모델들은 Long Short-Term Memory(LSTM) RNN, LSTM Recurrent CRF(LSTM-CRF), Bidirectional LSTM-CRF(Bi-LSTM-CRF) 등이 있다. LSTM-RNN은 기존 RNN의 그레디언트 소멸 문제(vanishing gradient problem)를 해결하였으며, LSTM-CRF는 출력 결과에 의존성을 부여하여 순차 데이터 문제에 더욱 최적화 하였다. Bi-LSTM-CRF는 과거입력자질과 미래입력자질을 함께 학습하는 방법으로 최근에 가장 좋은 성능을 보이고 있다. 이에 따라, 본 논문에서는 멘션탐지에 Bi-LSTM-CRF를 적용할 것을 제안하며, 각 딥 러닝 모델들에 대한 비교실험을 보인다.
PDF

Shallow Parsing on Grammatical Relations in Korean Sentences (한국어 문법관계에 대한 부분구문 분석)

Lee, Song-Wook;Seo, Jung-Yun
- Journal of KIISE:Software and Applications
- /
- v.32 no.10
- /
- pp.984-989
- /
- 2005
This study aims to identify grammatical relations (GRs) in Korean sentences. The key task is to find the GRs in sentences in terms of such GR categories as subject, object, and adverbial. To overcome this problem, we are fared with the many ambiguities. We propose a statistical model, which resolves the grammatical relational ambiguity first, and then finds correct noun phrases (NPs) arguments of given verb phrases (VP) by using the probabilities of the GRs given NPs and VPs in sentences. The proposed model uses the characteristics of the Korean language such as distance, no-crossing and case property. We attempt to estimate the probabilities of GR given an NP and a VP with Support Vector Machines (SVM) classifiers. Through an experiment with a tree and GR tagged corpus for training the model, we achieved an overall accuracy of $84.8\%,\;94.1\%,\;and\;84.8\%$ in identifying subject, object, and adverbial relations in sentences, respectively.
PDF KSCI

Syntactic Attraction of Subject-Verb Agreement (주어-동사 일치의 통사적 유인)

Jang, Soyeong;Kim, Yangsoon
- The Journal of the Convergence on Culture Technology
- /
- v.7 no.3
- /
- pp.353-358
- /
- 2021
This study provides the syntactic analysis for the agreement attraction by proposing three types of syntactic subject-verb agreement. Because subject-verb number agreement codifies the link between a predicate and its subject, it must be the purely syntactic processes of the head-to-head agreement or the feature percolation, where relevant agreement features percolate upward or downward through the hierarchical syntactic structure. The agreement errors are not affected by linear proximity or minimal interference, but instead are affected by the hierarchical relationship between an agreement target and a local attractor. The data in this paper includes the complex noun phrases with a modifier PP or a relative clause CP. Here, the [+PL] feature is suggested to be a local attractor for subject-verb agreement errors as a strong feature. Therefore, speakers tend to erroneously produce plural agreement for a singular subject in a main clause due to a plural NP in a modifier PP or plural agreement for a singular subject in a relative clause due to plural main subject.
https://doi.org/10.17703/JCCT.2021.7.3.353 인용 PDF KSCI

Chunking Using Automatic Constructed Syntactic Pattern Dictionary and Rule (자동 구축된 구문패턴사전과 규칙을 이용한 구묶음)

Im, Ji-Hui;Choe, Ho-Seop;Lee, Jung-Chul;Ock, Cheul-Young
- Annual Conference on Human and Language Technology
- /
- 2004.10d
- /
- pp.35-39
- /
- 2004
본 논문은 실용적인 구문분석기의 전단계로서, 자동 구축된 구문패턴사전과 규칙을 이용하여 구묶음하는 방법을 제안한다. 우선 규칙은 구문분석 말뭉치(30,875어절)를 대상으로 자동 추출된 고빈도의 규칙(Rewriting Rule)을 본 논문에 맞게 수동으로 구축하였다. 규칙은 조건부, 행위부로 이루어진 이진 규칙(binary rule)의 형태를 이루며, 명사구(NP), 수식어구(AP, DP), 인용구(X), 용언구(VP, VC)을 대상으로 15개를 구축하였다. 그리고 구문패턴은 중심어와 중심어 선행 요소의 특성뿐만 아니라 중심어 후행 요소도 고려하여 형식화시킨 것으로, 중심어의 복합용언 여부에 따라 일반용언패턴과 본+보조용언패턴으로 구분한다. 부분적인 언어 현상의 처리보다는 실세계에서 사용되는 수많은 문장들에 내재되어 있는 매우 광범위한 언어 현상의 처리를 하기 위해, 구문패턴은 형태소주석 말뭉치(460만 어절)을 대상으로 자동 구축하였다. 구축된 구문패턴사전과 규칙을 이용하여 구묶음을 수행한 결과 정확율 83.09%가 나타났다.
PDF

An Index System using Restrictive Distance (거리 제한을 이용한 색인 시스템)

Park, Chan-Ee;Kim, Sang-Bok
- Journal of the Korea Society of Computer and Information
- /
- v.11 no.1 s.39
- /
- pp.273-282
- /
- 2006
In this paper, we propose index method introducing distance concept in word by a method weighting word. This index method is frequent representing an inquiry word and document index and compound noun or more than two adjoin nouns or noun phrase, the farther the distance between these nouns, the fewer selected ratio decreases in index point is the aiming, this choose guide word candidate by existent weight grant method and distance between candidates chose candidate finally in index within 3 sentences. Using in these way I document of 100 kinds of newspaper, scientific treatise, web document and so on, showed the correctness rate resulted of newspaper 92.03% scientific treatise 95% web document 73.33%.
PDF

Search Result 150, Processing Time 0.027 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)