Search | Korea Science

Building an Annotated English-Vietnamese Parallel Corpus for Training Vietnamese-related NLPs

Dien Dinh;Kiem Hoang
- Proceedings of the IEEK Conference
- /
- summer
- /
- pp.103-109
- /
- 2004
In NLP (Natural Language Processing) tasks, the highest difficulty which computers had to face with, is the built-in ambiguity of Natural Languages. To disambiguate it, formerly, they based on human-devised rules. Building such a complete rule-set is time-consuming and labor-intensive task whilst it doesn't cover all the cases. Besides, when the scale of system increases, it is very difficult to control that rule-set. So, recently, many NLP tasks have changed from rule-based approaches into corpus-based approaches with large annotated corpora. Corpus-based NLP tasks for such popular languages as English, French, etc. have been well studied with satisfactory achievements. In contrast, corpus-based NLP tasks for Vietnamese are at a deadlock due to absence of annotated training data. Furthermore, hand-annotation of even reasonably well-determined features such as part-of-speech (POS) tags has proved to be labor intensive and costly. In this paper, we present our building an annotated English-Vietnamese parallel aligned corpus named EVC to train for Vietnamese-related NLP tasks such as Word Segmentation, POS-tagger, Word Order transfer, Word Sense Disambiguation, English-to-Vietnamese Machine Translation, etc.
PDF

A Korean Part-of-Speech Tagger using Simplified Eojeol-based unit (단순화된 어절을 단위로 하는 한국어 품사 태거)

Lee, Eui-Hyeon;Kim, Young-Gil;Shin, Jaehun;Kwon, Hong-Seok;Lee, Jong-Hyeok
- Annual Conference on Human and Language Technology
- /
- 2016.10a
- /
- pp.268-272
- /
- 2016
영어권 언어가 어절 단위로 품사를 부여하는 반면, 한국어는 굴절이 많이 일어나는 교착어로서 데이터부족 문제를 피하기 위해 형태소 단위로 품사를 부여한다. 이러한 구조적 차이 안에서 한국어에 적합한 품사 태깅 단위는 지속적으로 논의되어 왔으며 지금까지 음절, 형태소, 어절, 구가 제안되었다. 본 연구는 어절 단위로 태깅함으로써 야기되는 복잡한 품사 태그와 데이터부족 문제를 해소하기 위해 어절에서 주요 실질 형태소와 주요 형식 형태소만을 뽑아 새로운 어절을 생성하고, 생성된 단순한 어절에 대해 CRF 태깅을 수행하였다. 실험결과 평가 말뭉치에서 미등록 어절 등장 비율은 9.22%에서 5.63%로 38.95% 감소시키고, 어절단위 정확도를 85.04%에서 90.81%로 6.79% 향상시켰다.
PDF

A Morph Analyzer For MATES/CK (중한 기계 번역 시스템을 위한 형태소 분석기)

Kang, Won-Seok;Kim, Ji-Hyoun;Song, Young-Mi;Song, Hee-Jung;Huang, Jin-Xia;Chae, Young-Soog;Choi, Key-Sun
- Annual Conference on Human and Language Technology
- /
- 2000.10d
- /
- pp.331-336
- /
- 2000
MATES/CK는 기계번역 시스템에서 전통적으로 사용하고 있는 세 단계(분석/변환/생성)에 의해서 중한 번역을 수행하는 시스템이다. MATES/CK는 시스템 성능을 높이기 위해 패턴 기반과 통계적 정보를 이용한다. 태거(Tagger)는 중국어 단어 분리를 최장일치법으로 수행하기 때문에 일부 단어에 대해 오류를 범하게 되고 품사(POS : Part Of Speech) 태깅 시 확률적 정보만 이용하여 특정 단어가 다 품사인 경우 그 단어에 대해 특정 품사만 태깅되는 문제점이 발생한다. 또한 중국어 및 외국어 인명 및 지명에 대한 미등록들에 대해서도 올바른 결과를 도출하지 못한다. 사전에 있어서 텍스트 기반으로 존재하여 이를 관리하기에 힘이 든다. 본 논문에서는 단어 분리 오류 및 품사 태깅 오류를 해결하기 위해 중국어 태깅 제약 규칙을 적용하는 방법을 제시하고 중국어 및 외국어 인명/지명에 대한 미등록어 처리방법을 제시한다. 또한 중국어 사전 관리에 대해 알아본다.
PDF

Application portable Part-Of-Speech tagger mapping (응용을 위한 품사 태깅 시스템의 매핑)

Kim, Jun-Seok;Cha, Jung-Won;Lee, Geun-Bae
- Annual Conference on Human and Language Technology
- /
- 2000.10d
- /
- pp.368-375
- /
- 2000
품사 태깅 시스템은 자연 언어 처리의 가장 기본이 되는 부분으로 상위 자연 언어 처리 분야인 구문분석, 의미분석의 전처리로 사용되거나, 기계번역, 정보검색이나 음성인식 및 합성 등과 같은 많은 응용 시스템을 위해서도 필요하다. 이렇게 여러 가지 목적을 위해 품사 태깅 시스템은 존재하는데, 각각의 응용을 위해서 최적화된 태깅 시스템을 따로 구성하기도 하고, 하나의 태깅 시스템을 여러 가지 응용을 위해서 사용하기도 한다. 이때, 문제가 되는 것 중에 하나는 각 응용마다 요구하는 품사 태그 세트가 다르다는 것이다. 품사 태그세트가 고정되어 있다면 어떤 응용을 위해서는 사용되는 품사 태그세트가 너무 적어서 문제가 되고, 반대로 품사태그세트가 너무 많아서 시스템의 수행속도가 중요시되는 응용에서 성능저하의 요인이 되기도 한다. 본 논문에서는 하나의 태깅 시스템의 품사태그세트를 조절할 수 있도록 하여 몇 가지 응용시스템에 맞게 최적화시킬 수 있는 방법론을 제시하고 실험을 통해서 시스템의 성능, 유지보수 및 시스템의 여러 리소스 관리 측면에서도 가장 효율적인 방법론임을 입증하고자 한다.
PDF

Implementation of A Morphological Analyzer Based on Pseudo-morpheme for Large Vocabulary Speech Recognizing (대어휘 음성인식을 위한 의사형태소 분석 시스템의 구현)

양승원
- Journal of Korea Society of Industrial Information Systems
- /
- v.4 no.2
- /
- pp.102-108
- /
- 1999
It is important to decide processing unit in the large vocabulary speech recognition system we propose a Pseudo-Morpheme as the recognition unit to resolve the problems in the recognition systems using the phrase or the general morpheme. We implement a morphological analysis system and tagger for Pseudo-Morpheme. The speech processing system using this pseudo-morpheme can get better result than other systems using the phrase or the general morpheme. So, the quality of the whole spoken language translation system can be improved. The analysis-ratio of our implemented system is similar to the common morphological analysis systems.
PDF

PharmacoNER Tagger: a deep learning-based tool for automatically finding chemicals and drugs in Spanish medical texts

Armengol-Estape, Jordi;Soares, Felipe;Marimon, Montserrat;Krallinger, Martin
- Genomics & Informatics
- /
- v.17 no.2
- /
- pp.15.1-15.7
- /
- 2019
Automatically detecting mentions of pharmaceutical drugs and chemical substances is key for the subsequent extraction of relations of chemicals with other biomedical entities such as genes, proteins, diseases, adverse reactions or symptoms. The identification of drug mentions is also a prior step for complex event types such as drug dosage recognition, duration of medical treatments or drug repurposing. Formally, this task is known as named entity recognition (NER), meaning automatically identifying mentions of predefined entities of interest in running text. In the domain of medical texts, for chemical entity recognition (CER), techniques based on hand-crafted rules and graph-based models can provide adequate performance. In the recent years, the field of natural language processing has mainly pivoted to deep learning and state-of-the-art results for most tasks involving natural language are usually obtained with artificial neural networks. Competitive resources for drug name recognition in English medical texts are already available and heavily used, while for other languages such as Spanish these tools, although clearly needed were missing. In this work, we adapt an existing neural NER system, NeuroNER, to the particular domain of Spanish clinical case texts, and extend the neural network to be able to take into account additional features apart from the plain text. NeuroNER can be considered a competitive baseline system for Spanish drug and CER promoted by the Spanish national plan for the advancement of language technologies (Plan TL).
https://doi.org/10.5808/GI.2019.17.2.e15 인용 PDF KSCI

A Corpus-Based Longitudinal Study of Diction in Chinese and British News Reports on Chang'e Project

Lu, Rong;Xie, Xue;Qi, Jiashuang;Ali, Afida Mohamad;Zhao, Jie
- Asia Pacific Journal of Corpus Research
- /
- v.3 no.1
- /
- pp.1-20
- /
- 2022
As a milestone progression in China's space exploration history, Chang'e Project has attracted a lot of media attention since its first launching. This study aims to examine and compare the similarities and differences between the Chinese media and the British media in using nouns, verbs, and adjectives to report the Chang'e Project. After categorising the documents based on specific project phases, we created two diachronic corpora to explore the linguistic shifts and similarities and differences of diction employed by the Chinese and British media on the Chang'e Project ideology. This longitudinal study was performed with Lancsbox and the CLAWS web tagger through critical discourse analysis as the theoretical framework. The findings of the current study showed that the Chang'e Project coverage in both media increased on an annual basis, especially after 2019. In contrast to the objectivity and positivity in the Chinese Media, the British Media seemed to be more subjective with more appraisal adjectives in the news reports. Nonetheless, both countries were trying to be objective and formal in choosing nouns and verbs. Ideology-wise, the Chinese news media reports portrayed more positivity on domestic circumstances while the British counterpart was typically more critical. Notably, the study outcomes could catalyse future research on the Chang'e Project and facilitate diplomatic policies.
https://doi.org/10.22925/apjcr.2022.3.1.1 인용 PDF KSCI

A Study on the Application of LibraryThing Folksonomy Tags through the Analysis of Elements related with Work (저작관련 요소분석을 통한 폭소노미 태그의 활용 방안에 관한 연구: LibraryThing을 중심으로)

Kim, Dong-Suk;Chung, Yeon-Kyoung
- Journal of the Korean Society for information Management
- /
- v.27 no.1
- /
- pp.41-60
- /
- 2010
This study aims to analyze the properties of the tags used in the fiction genre, the structural aspect of the patterns and the contents of the tags by utilizing LibraryThing, where the tags are assigned in work units of FRBR. A comparative analysis was conducted in terms of the level of association between the descriptive terms in bibliography and LCSH terms. The study also examined the sources of the tags not included in the bibliographic descriptions or LCSHs, what aspects of work they represented, and the terms used as tags in relation to the work. By restricting the study to a single genre, a number of tags that reflected the characteristics of fiction (three elements of the fiction which are theme, plot, style and three elements of the fiction composition which are character, event, setting) were extracted. This study finds out the role of the tag making up the taxonomy and proposes a new direction for the tagging system by demonstrating the possibility of using tags as facets in information organization and retrieval.
https://doi.org/10.3743/KOSIM.2010.27.1.041 인용 PDF

KONG-DB: Korean Novel Geo-name DB & Search and Visualization System Using Dictionary from the Web (KONG-DB: 웹 상의 어휘 사전을 활용한 한국 소설 지명 DB, 검색 및 시각화 시스템)

Park, Sung Hee
- Journal of the Korean Society for information Management
- /
- v.33 no.3
- /
- pp.321-343
- /
- 2016
This study aimed to design a semi-automatic web-based pilot system 1) to build a Korean novel geo-name, 2) to update the database using automatic geo-name extraction for a scalable database, and 3) to retrieve/visualize the usage of an old geo-name on the map. In particular, the problem of extracting novel geo-names, which are currently obsolete, is difficult to solve because obtaining a corpus used for training dataset is burden. To build a corpus for training data, an admin tool, HTML crawler and parser in Python, crawled geo-names and usages from a vocabulary dictionary for Korean New Novel enough to train a named entity tagger for extracting even novel geo-names not shown up in a training corpus. By means of a training corpus and an automatic extraction tool, the geo-name database was made scalable. In addition, the system can visualize the geo-name on the map. The work of study also designed, implemented the prototype and empirically verified the validity of the pilot system. Lastly, items to be improved have also been addressed.
https://doi.org/10.3743/KOSIM.2016.33.3.321 인용 PDF KSCI

An Experimental Study on Opinion Classification Using Supervised Latent Semantic Indexing(LSI) (지도적 잠재의미색인(LSI)기법을 이용한 의견 문서 자동 분류에 관한 실험적 연구)

Lee, Ji-Hye;Chung, Young-Mee
- Journal of the Korean Society for information Management
- /
- v.26 no.3
- /
- pp.451-462
- /
- 2009
The aim of this study is to apply latent semantic indexing(LSI) techniques for efficient automatic classification of opinionated documents. For the experiments, we collected 1,000 opinionated documents such as reviews and news, with 500 among them labelled as positive documents and the remaining 500 as negative. In this study, sets of content words and sentiment words were extracted using a POS tagger in order to identify the optimal feature set in opinion classification. Findings addressed that it was more effective to employ LSI techniques than using a term indexing method in sentiment classification. The best performance was achieved by a supervised LSI technique.
https://doi.org/10.3743/KOSIM.2009.26.3.451 인용 PDF

Search Result 62, Processing Time 0.021 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)