• Title/Summary/Keyword: 언어 자료

Search Result 1,224, Processing Time 0.026 seconds

Construction of a Parallel Corpus for Instant Messenger Spelling Correction and Related Issues (메신저 맞춤법 교정 병렬 말뭉치의 구축과 쟁점)

  • HUANG YINXIA;Jin-san An;Kil-im Nam
    • Annual Conference on Human and Language Technology
    • /
    • 2022.10a
    • /
    • pp.545-550
    • /
    • 2022
  • 본 연구의 목적은 2021년 메신저 언어 200만 어절을 대상으로 수행된 맞춤법 교정 병렬 말뭉치의 설계와 구축의 쟁점을 소개하고, 교정 말뭉치의 주요 교정 및 주석 내용을 기술함으로써 맞춤법 교정 병렬 말뭉치의 특성을 분석하는 것이다. 2021년 맞춤법 교정 병렬 말뭉치의 주요 목표는 메신저 언어의 특수성을 살림과 동시에 형태소 분석이나 기계 번역 등 한국어 처리 도구가 분석할 수 있는 수준으로 교정하는 다소 상충되는 목적을 구현하는 것이었는데, 이는 교정의 수준과 병렬의 단위 설정 등 상당한 쟁점을 내포한다. 본 연구에서는 말뭉치 구축 시점에서 미처 논의하지 못한 교정 수준의 쟁점과 교정 전후의 통계적 특성을 함께 논의하고자 하며, 다음과 같은 몇 가지 하위 내용을 중심으로 논의하고자 한다.첫째, 맞춤법 교정 병렬 말뭉치의 구조 설계와 구축 절차에 대한 논의로, 2022년 초 국내 최초로 공개된 한국어 맞춤법 교정 병렬 말뭉치('모두의 말뭉치'의 일부)의 구축 과정에서 논의되어 온 말뭉치 구조 설계와 구축 절차를 논의한다. 둘째, 문장 단위로 정렬된 맞춤법 교정 말뭉치에서 관찰 가능한 띄어쓰기, 미등재어, 부호형 이모티콘 등의 메신저 언어의 몇 가지 특성을 살펴본다. 마지막으로, 2021년 메신저 맞춤법 교정 말뭉치의 구축 단계에서 미처 논의되지 못한 남은 문제들을 각각 데이터 구조 설계와 구축 차원의 주요 쟁점을 중심으로 논의한다. 특히 메신저 맞춤법 병렬 말뭉치의 주요 목표인 사전학습 언어모델의 학습데이터로서의 가치와 메신저 언어 연구의 기반 자료 구축의 관점에서 맞춤법 교정 병렬 말뭉치 구축의 의의와 향후 과제를 논의하고자 한다.

  • PDF

A Study on the Connecting Method of Query and Legal Cases Using Doc2Vec Document Embedding (Doc2Vec 문서 임베딩을 이용한 질의문과 판례 자동 연결 방안 연구)

  • Kang, Ye-Jee;Kang, Hye-Rin;Park, Seo-Yoon;Jang, Yeon-Ji;Kim, Han-Saem
    • Annual Conference on Human and Language Technology
    • /
    • 2020.10a
    • /
    • pp.76-81
    • /
    • 2020
  • 법률 전문 지식이 없는 사람들이 법률 정보 검색을 성공적으로 하기 위해서는 일반 용어를 검색하더라도 전문 용어가 사용된 법령정보가 검색되어야 한다. 하지만 현 판례 검색 시스템은 사용자 선호도 검색이 불가능하며, 일반 용어를 사용하여 검색하면 사용자가 원하는 전문 자료를 도출하는 데 어려움이 있다. 이에 본 논문에서는 일반용어가 사용된 질의문과 전문용어가 사용된 판례를 자동으로 연결해 주고자 하였다. 질의문과 연관된 판례를 자동으로 연결해 주기 위해 전문용어가 사용된 전문가 답변을 바탕으로 문서분류에 높은 성능을 보이는 Doc2Vec을 이용한다. Doc2Vec 문서 임베딩 기법을 이용하여 전문용어가 사용된 전문가 답변과 유사한 답변을 제안하여 비슷한 주제의 답변들끼리 분류하였다. 또한 전문가 답변과 유사도가 높은 판례를 제안하여 질의문에 해당하는 판례를 자동으로 연결하였다.

  • PDF

Identifying Optimum Features for Abbreviation Disambiguation in Biomedical Domain (생의학 도메인에서 약어 중의성 해결을 위한 최적 자질의 규명)

  • Lim, Ho-Gun;Seo, Hee-Cheol;Kim, Seon-Ho;Rim, Hae-Chang
    • Annual Conference on Human and Language Technology
    • /
    • 2004.10d
    • /
    • pp.173-180
    • /
    • 2004
  • 생의학 도메인에서 약어 중의성 해결이란 생의학 문서에 나타난 약어의 원래 형태(long form)를 판별하는 작업이다. 본 논문은 생의학 도메인에서 약어 중의성 해결에 적합한 자질들을 실험적으로 탐색하는데 목적이 있다. 이를 위해서 약어 중의성 해결에 사용할 문맥을 전역 문맥(topical context)과 지역 문맥(local context)으로 구분하고, 각각의 문맥에서 스테밍(stemming), 불용어 제거, 품사 부착 등의 과정을 통해서 다양한 자질들을 고려하도록 한다. 생의학 도메인에서 약어 중의성 해결을 위한 실험 자료의 부족을 해결하기 위해서, 학습 자료와 평가 자료를 자동으로 구축했으며, 평가를 위한 약어로는 기존 연구에서 사용된 두 가지 약어 목록을 사용했다. 또한 단순 베이지언 모델(Naive Bayesian Model)을 이용해서 각 자질들의 유용성을 평가하였다 실험 결과, 전역 문맥이 지역 문맥보다 더 좋은 성능을 보였으며, 전역 문맥에서는 불용어만을 제거한 경우가 각각의 평가 자료에서 94.2%와 96.2%로 가장 좋은 결과를 보였으며, 전역 문맥과 지역 문맥을 함께 사용하는 경우에 각각의 평가 자료에서 1.8%와 0.3%의 성능 향상이 있었다.

  • PDF

The Help of Experienced Dental Hygienists Turnover Verbal Abuse and Emotional Reaction, and the Resulting Relationship (치과위생사가 경험하는 언어폭력과 그에 따른 정서적 반응 및 이직의도와의 관계)

  • Lee, Jung-Hwa;Choi, Jung-Mi;Lee, Yeong-Ae
    • Journal of dental hygiene science
    • /
    • v.14 no.4
    • /
    • pp.563-570
    • /
    • 2014
  • The purpose of this study was to investigate the degree of verbal violence against dental hygienists, their emotional reaction and the relation between their intention of job transfer and verbal violence so that it could offer the basic data for developing the way how to cope with verbal violence and for improving their performance. Two hundred fifty-seven dental hygienists working for dentists' in Busan were interviewed from May 17 to 31, 2014 to collect data, of which analysis was as follows: 1) As a result of verbal violence done by patients and their guardians, 80.5% said that they experienced crude language with 17.5% forceful and imperative sentence, and 13.2% ignorant statements about their job. They were exposed to verbal violence once or twice every 6 months. As a result of researching verbal violence of co-working senior or junior hygienists, 52.1% answered that their co-working senior or junior hygienists talked crude language and 38.1% said their co-workers happened to say crude language to them. The crude language experience was relatively high as 20.2% and once or twice a week. As a result of verbal violence done by dentists, 47.5% said that they've heard crude language and 34.6% said that they experienced forceful and imperative sentence. 2) The overall average of the intention to transfer their job was $3.06{\pm}1.03$, while the highest intention of job transfer was $3.11{\pm}0.91$ where they said I have once wanted to transfer my job. 3) As a result of seeing the relation among verbal violence, emotional reaction and the intention of job transfer, there was co-relation between verbal violence and the patients' age (p<0.01); there were also co-relation between verbal violence of patients, co-workers and dentists (p<0.01). There was also significant relation between emotional reaction on verbal violence and their intention of job transfer (p<0.001).

The Relationship Among Domain-General Creativity, Linguistic Intelligence, Korean Language Grade and Linguistic Creativity of Elementary School Student (초등학생의 일반창의성, 언어지능, 국어성적과 언어창의성 간의 관계연구)

  • Park, Jung-Hwan;Hong, Mi-Sun;Lew, Kyoung-Hoon
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.14 no.8
    • /
    • pp.3760-3767
    • /
    • 2013
  • The purpose of this study is to investigate the relationship among domain-general creativity, linguistic intelligence, Korean language grade and linguistic creativity of elementary school student. And to confirm the relative predictive power of domain-general creativity variables in predicting elementary school students' linguistic creativity. The instruments used in this study were 'TTCT', 'Essay writing' and 'Linguistic intelligence ' and school grade of Korean language. Self-reported response data on these instruments from 338, 4th grade elementary school students in Seoul were analyzed. The data were analyzed with descriptive statistics, Pearson correlations, multiple stepwise regression analysis and ANOVA by using SPSS 18.0. The major results of this study were as follows; First, the correlations among domain-general creativity, Korean language grade and linguistic creativity were significant. Second, Abstractness of title were the best predictor of linguistic creativity in elementary school students.

A Content Analysis of Journal Articles Using the Language Network Analysis Methods (언어 네트워크 분석 방법을 활용한 학술논문의 내용분석)

  • Lee, Soo-Sang
    • Journal of the Korean Society for information Management
    • /
    • v.31 no.4
    • /
    • pp.49-68
    • /
    • 2014
  • The purpose of this study is to perform content analysis of research articles using the language network analysis method in Korea and catch the basic point of the language network analysis method. Six analytical categories are used for content analysis: types of language text, methods of keyword selection, methods of forming co-occurrence relation, methods of constructing network, network analytic tools and indexes. From the results of content analysis, this study found out various features as follows. The major types of language text are research articles and interview texts. The keywords were selected from words which are extracted from text content. To form co-occurrence relation between keywords, there use the co-occurrence count. The constructed networks are multiple-type networks rather than single-type ones. The network analytic tools such as NetMiner, UCINET/NetDraw, NodeXL, Pajek are used. The major analytic indexes are including density, centralities, sub-networks, etc. These features can be used to form the basis of the language network analysis method.

An Analysis of Language Activity Contents for Young Children from the Nuri Curriculum Teacher's Guidebooks for Age 3-5 (3~5세 누리과정 교사용 지도서에 나타난 유아 언어교육 활동 내용 분석)

  • Han, Sun-Ah;Kwak, Jung-In
    • The Journal of the Korea Contents Association
    • /
    • v.13 no.7
    • /
    • pp.511-521
    • /
    • 2013
  • The purpose of this study is to review the perspective on early childhood language education by analyzing language activities specified in the Teacher's guide to Nuri Curriculum for Children between Age 3 to 5. In the pursuit of this purpose, 966 language educational activities suggested in 32 guidebooks(10 for age 3, 11 for age 4, 11 for age 5 - divided by life themes) have been chosen as the analysis object and analyzed based on the following category; subordinate scope, and activity type. This analysis showed that children aged 3~5 start their language activities in the order of talking, listening, reading and writing (under the subordinate scope category), and favors activities in the order of 'fairy tale/poem', 'story telling' and 'verbal section'. In conclusion, it has been proven that each category is concentrating on 1~2 activities and the proportion varies depending on the age. Based on the above result, we intend to examine the current situation of language education and use this study as the preliminary data to provide a proper direction for early childhood language education.

Development of An Intelligent Agent Shell Supporting An Integrated Agent Building Language (통합 에이전트 구축 언어를 지원하는 지능형 에이전트 쉘의 개발)

  • Chang, Hai-Jin
    • The Transactions of the Korea Information Processing Society
    • /
    • v.6 no.12
    • /
    • pp.3548-3558
    • /
    • 1999
  • There are many kinds of multi-agent frameworks which support the high-level knowledge representation languages for providing intelligence to their agents. But, the agent programming interfaces of the frameworks require to use some general-purpose programming languages as well as tile knowledge representation languages. In general, knowledge representation languages and general-purpose programming languages are different in their levels and data representation models. The differences can make the problems about tile coupling of the elements which are necessary for developing intelligent agents. This paper describes a new type of intelligent agent shell INAS(INtelligent Agent Shell) version 2 which has developed to cope with the problems. Unlike the previous agent frameworks, INAS supports a high-level integrated agent building language for building intelligent agents by itself. Therefore, the development of intelligent agents by using INAS version 2 does not suffer from the problems of the previous agent frameworks. Through the development of several intelligent agents, we experienced that the agent building language of INAS version 2 could reduce the difficulties of developing intelligent agents.

  • PDF

A Case Study for Migration from SGML Document to XML Documents (SGML 문서를 XML 문서로 변환하는 사례 연구)

  • Cho, Min-Ho;Ryew, Sung-Yul;Park, Si-Hyoung
    • Journal of KIISE:Computing Practices and Letters
    • /
    • v.7 no.6
    • /
    • pp.653-660
    • /
    • 2001
  • Recently, The range of Internet based information environment is spreading over core business area, as well as simple information provision area. Especially, with spreading WWW technology, markup language based technology is emerging as an important part in Internet based business. But, the data made by SGML can only see by using SGML Browser, so it has some problem in information providing at Internet, and compatibility of data between Data source. So, this study suggests essential architecture and technique for migrating from SGML to XML environment. In our study, we use 600MB SGML data that are selected from 3Tera DataBase of SGML as testing target for migration. We can reduce data displaying time after migration, can do mobile computing which is based on Internet as a result of this study. And the same technique and idea that is used in this study can apply to more large SGML Environment without changing. So, It will be very helpful to the reader who is interesting to migrate from SGML doc to XML doc.

  • PDF