• Title/Summary/Keyword: 서울 코퍼스

Search Result 16, Processing Time 0.023 seconds

Study on Personification of Korean open domain Dialog system: Focusing on honorific expression under changes of social variations (한국어 오픈도메인 대화 시스템의 의인화 연구: 사회적 변인에 따른 상대높임법 중심)

  • Choi, Nam-Kyu;Min, Byeong-Cheol;Cho, Woo-Ri;Min, Kyung-eun;Jeong, Han-kyeol;Uprety, Sudan Prasad
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2022.11a
    • /
    • pp.393-395
    • /
    • 2022
  • 실제 대화에서는 다양한 화자와 청자간의 사회적 위치와 관계 등의 사회적 변인에 따라 다양한 상대높임법이 존재한다. 제안하는 상대높임법 중심의 대화시스템 아키텍처를 설명하기에 앞서 배경지식 및 관련연구로 규칙/코퍼스 기반 대화시스템을 소개하고, 상대높임법을 포함하는 공손법처리에 대한 기존 연구들의 제약사항을 논의한다. 본 연구에서는 한국어 상대높임법을 정의 및 사회적 변인 모델링하고 이를 구현하기 위한 대화시스템 아키텍처 방안을 제안한다.

Phoneme distribution and phonological processes of orthographic and pronounced phrasal words in light of syllable structure in the Seoul Corpus (음절구조로 본 서울코퍼스의 글 어절과 말 어절의 음소분포와 음운변동)

  • Yang, Byunggon
    • Phonetics and Speech Sciences
    • /
    • v.8 no.3
    • /
    • pp.1-9
    • /
    • 2016
  • This paper investigated the phoneme distribution and phonological processes of orthographic and pronounced phrasal words in light of syllable structure in the Seoul Corpus in order to provide linguists and phoneticians with a clearer understanding of the Korean language system. To achieve the goal, the phrasal words were extracted from the transcribed label scripts of the Seoul Corpus using Praat. Following this, the onsets, peaks, codas and syllable types of the phrasal words were analyzed using an R script. Results revealed that k0 was most frequently used as an onset in both orthographic and pronounced phrasal words. Also, aa was the most favored vowel in the Korean syllable peak with fewer phonological processes in its pronounced form. The total proportion of all diphthongs according to the frequency of the peaks in the orthographic phrasal words was 8.8%, which was almost double those found in the pronounced phrasal words. For the codas, nn accounted for 34.4% of the total pronounced phrasal words and was the varied form. From syllable type classification of the Corpus, CV appeared to be the most frequent type followed by CVC, V, and VC from the orthographic forms. Overall, the onsets were more prevalent in the pronunciation more than the codas. From the results, this paper concluded that an analysis of phoneme distribution and phonological processes in light of syllable structure can contribute greatly to the understanding of the phonology of spoken Korean.

Phonological processes of vowels in pronounced phrasal words of the Seoul Corpus by gender and age groups (서울코퍼스의 성별·연령 집단별 말 어절 모음에 나타난 음운변동)

  • Yang, Byunggon
    • Phonetics and Speech Sciences
    • /
    • v.9 no.2
    • /
    • pp.23-29
    • /
    • 2017
  • This paper investigated the phonological processes of monophthongs and diphthongs in pronounced phrasal words of the Seoul Corpus by gender and age groups in order to provide linguists and phoneticians with a clearer understanding of the spoken Korean. Both orthographic and pronounced phrasal words were extracted from the transcribed label scripts of the Corpus using Praat. Then, phonological processes of monophthongs and diphthongs were tabulated using an R script after syllabifying the phrasal words into separate components. Results revealed that 97% of the number of syllables in the orthographic and pronounced phrasal words were the same while 65.8% showed difference in the syllable structure. 90.5% of the vowels in the orthographic phrasal words were realized in the pronounced phrasal words. A Chi-square test of independence was performed to obtain a significant dependence in the distribution of phonological process types of male and female groups along with a very strong correlation. Female group changed the diphthong yo into yv at the end of the pronounced phrasal words more often than the male group did. Age groups also showed a significant dependence in the distribution of phonological process types along with a very strong correlation. Females in the 40s produced the diphthong yv and made the vowel raising at the end of the pronounced phrasal words most often among the gender and age groups. From the results, this paper concludes that an analysis of phonological processes in light of syllable structure can contribute greatly to the understanding of the spoken Korean.

CRNN-Based Korean Phoneme Recognition Model with CTC Algorithm (CTC를 적용한 CRNN 기반 한국어 음소인식 모델 연구)

  • Hong, Yoonseok;Ki, Kyungseo;Gweon, Gahgene
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.8 no.3
    • /
    • pp.115-122
    • /
    • 2019
  • For Korean phoneme recognition, Hidden Markov-Gaussian Mixture model(HMM-GMM) or hybrid models which combine artificial neural network with HMM have been mainly used. However, current approach has limitations in that such models require force-aligned corpus training data that is manually annotated by experts. Recently, researchers used neural network based phoneme recognition model which combines recurrent neural network(RNN)-based structure with connectionist temporal classification(CTC) algorithm to overcome the problem of obtaining manually annotated training data. Yet, in terms of implementation, these RNN-based models have another difficulty in that the amount of data gets larger as the structure gets more sophisticated. This problem of large data size is particularly problematic in the Korean language, which lacks refined corpora. In this study, we introduce CTC algorithm that does not require force-alignment to create a Korean phoneme recognition model. Specifically, the phoneme recognition model is based on convolutional neural network(CNN) which requires relatively small amount of data and can be trained faster when compared to RNN based models. We present the results from two different experiments and a resulting best performing phoneme recognition model which distinguishes 49 Korean phonemes. The best performing phoneme recognition model combines CNN with 3hop Bidirectional LSTM with the final Phoneme Error Rate(PER) at 3.26. The PER is a considerable improvement compared to existing Korean phoneme recognition models that report PER ranging from 10 to 12.

A Genre Analysis of Newspaper Articles for Korean Language Education -Based on the linguistic analysis of newspaper articles and reading materials in Korean language textbooks- (한국어 읽기 교육을 위한 기사문 장르분석 -신문기사 및 교재 기사문의 언어학적 분석을 바탕으로-)

  • Lee, Seungyeon;Sim, Jiyeon;Shin, Jungha
    • Journal of Korean language education
    • /
    • v.28 no.3
    • /
    • pp.53-83
    • /
    • 2017
  • The goal of this study is to examine whether the genre characteristics of newspaper articles are appropriately reflected in Korean language textbooks. For the purpose of this study, two corpora were built with 17 textbook articles and 60 newspaper articles respectively. The average sentence length and frequency of vocabulary in each corpus were measured. It was found that the sentences of articles in textbooks tended to have longer sentence length and more complicated structures than the articles in newspapers. For instance, sentences in the textbook articles had more verbal endings, such as conjunctive and transforming endings. On the other hand, in case of vocabulary representing 'timeliness', there was a high frequency of adverbs and nouns which were related to year, month, and time in actual articles, while it is found to be very limited in textbooks. Also, typical translative styles such as '-ko itta', '-e ttareumyun' were more prominent in textbooks than in newspaper articles. In the case of abbreviated and omitted form of particles, this was a characteristic that appeared only in actual articles because of the constraint of space. It is significant that this paper offers suggestions for the development of reading materials for Korean language education by revealing that the genre typology of actual newspaper articles is not adequately reflected in current textbooks.

Pronunciation of the Korean diphthong /jo/: Phonetic realizations and acoustic properties (한국어 /ㅛ/의 발음 양상 연구: 발음형 빈도와 음향적 특징을 중심으로)

  • Hyangwon Lee
    • Phonetics and Speech Sciences
    • /
    • v.15 no.1
    • /
    • pp.9-17
    • /
    • 2023
  • The purpose of this study is to determine how the Korean diphthong /jo/ shows phonetic variation in various linguistic environments. The pronunciation of /jo/ is discussed, focusing on the relationship between phonetic variation and the distribution range of vowels. The location in a word (monosyllable, word-initial, word-medial, word-final) and word class (content word, function word) were analyzed using the speech of 10 female speakers of the Seoul Corpus. As a result of determining the frequency of appearance of /jo/ in each environment, the pronunciation type and word class were affected by the location in a word. Frequent phonetic reduction was observed in the function word /jo/ in the acoustic analysis. The word class did not change the average phonetic values of /jo/, but changed the distribution of individual tokens. These results indicate that the linguistic environment affects the phonetic distribution of vowels.