• Title/Summary/Keyword: 코퍼스 분석

Search Result 206, Processing Time 0.021 seconds

The effects of speakers' age on temporal features of speech among healthy young, middle-aged, and older adults (연령세대에 따른 말 산출의 시간적 특성: 말속도와 쉼을 중심으로)

  • Kim, Yeji;Lee, Song-min;Choi, Min-kyung;Jung, Sang-min;Sung, Jee Eun;Lee, Youngmee
    • Phonetics and Speech Sciences
    • /
    • v.14 no.1
    • /
    • pp.37-47
    • /
    • 2022
  • The purpose of the this study is to observe the effects of healthy adults' age on temporal features of speech and identify which could differentiate older and young adults. We examined speech rates(i.e., overall speaking rate, articulation rate), occurrence of pause, and duration of pause per utterance by utilizing the National Institute of Korean Language's open corpus. We selected a total of 30 healthy adults (10 young, 10 middle-aged, and 10 older adults) in this study. There were significant differences among the groups in the overall speaking rate, articulation rate, total occurrence of pause, the occurrence of pause between syntactic words, total duration of pause, and duration of pause between syntactic words. The older and middle-aged adults showed slower speech rates and longer and more frequent pause than young adults. But there were no significant differences among the three groups in terms of pause within syntactic word. The overall speaking rate significantly differentiated older adults from young adults. These findings suggested that the effect of speakers' age was reflected in gradual changes in the temporal features of their speech.

The Ability of L2 LSTM Language Models to Learn the Filler-Gap Dependency

  • Kim, Euhee
    • Journal of the Korea Society of Computer and Information
    • /
    • v.25 no.11
    • /
    • pp.27-40
    • /
    • 2020
  • In this paper, we investigate the correlation between the amount of English sentences that Korean English learners (L2ers) are exposed to and their sentence processing patterns by examining what Long Short-Term Memory (LSTM) language models (LMs) can learn about implicit syntactic relationship: that is, the filler-gap dependency. The filler-gap dependency refers to a relationship between a (wh-)filler, which is a wh-phrase like 'what' or 'who' overtly in clause-peripheral position, and its gap in clause-internal position, which is an invisible, empty syntactic position to be filled by the (wh-)filler for proper interpretation. Here to implement L2ers' English learning, we build LSTM LMs that in turn learn a subset of the known restrictions on the filler-gap dependency from English sentences in the L2 corpus that L2ers can potentially encounter in their English learning. Examining LSTM LMs' behaviors on controlled sentences designed with the filler-gap dependency, we show the characteristics of L2ers' sentence processing using the information-theoretic metric of surprisal that quantifies violations of the filler-gap dependency or wh-licensing interaction effects. Furthermore, comparing L2ers' LMs with native speakers' LM in light of processing the filler-gap dependency, we not only note that in their sentence processing both L2ers' LM and native speakers' LM can track abstract syntactic structures involved in the filler-gap dependency, but also show using linear mixed-effects regression models that there exist significant differences between them in processing such a dependency.

A Study on the general language use of ROOJIN : in Headline Database of Newspaper Articles and Balanced Corpus of Contemporary Written Japanese by KOKKEN (현대일본문장어의 「노인(老人)」사용실태 - 国硏「ことばに関する新聞記事見出しデ?タベ?ス」 「現代日本語書き言葉均衡コ?パス」를 분석대상으로)

  • Oh, Mi sun
    • Cross-Cultural Studies
    • /
    • v.25
    • /
    • pp.627-648
    • /
    • 2011
  • The study analyzed a diachronic distribution, social meanings and social evaluations of ROOJIN. 'Headline Database of Newspaper Articles' and 'Balanced Corpus of Contemporary Written Japanese' by KOKKEN were used as research data. There were 305 newspaper articles (About 0.2%) which contained the word ROOJIN at 'Headline Database of Newspaper Articles'. The number of newspaper articles related to ROOJIN started to increase in a rapid rate in 1972 and 1973. They were also increased in 1976, from 1981 to 1987, 1992 and 1993. The reasons of increasing of newspaper articles related to ROOJIN on those 4 periods of time could be summarized as follows. Firstly, there was a increasement of ROOJIN who are lonely, are not able to move about freely or live alone. Secondly, the understanding of a symptom of aging called BOKE was necessary. Thirdly, there were negative evaluations in a society towards ROOJIN. There were 453 cases which contained the word ROOJIN at 'Balanced Corpus of Contemporary Written Japanese' on the data since 2000. The most frequently used words were ones that are related to senior care facilities. There were 109 cases (24%) which contain those words. '~SISETSU', '~SENTA-', '~HO-MU' were presented as words related to senior care facilities. Among them, 78 cases contained the word '~HO-MU' which was similar to a home with family members. The second most frequently used words were ones that are related to 'welfare for the aged' and they are led by 'medical care for the aged'. They occupied about 8%. Institutionalization of medical care for the aged, medical expenses, nursing were presented as words related to 'medical care for the aged'. Words that were related to 'welfare for the aged' led by 'senior care facilities' and 'medical care for the aged' occupied about 32% of research data. As mentioned above, problems of the aged in Modern Japan such as negative evaluations in a society towards ROOJIN, ROOJIN who are lonely, are not able to move about freely or live alone, BOKE could be identified by analyzing the data. Also, The frequent usage of words such as 'Home for the aged', 'medical care for the aged' and 'nursing' could be identified. The outcome of analysis suggested that a family traditionally had a function of solving problems of the aged but that function was reduced in modern Japan. It also suggested that there was a tendency to outsource problems of the aged as much as possible.

A realization of pauses in utterance across speech style, gender, and generation (과제, 성별, 세대에 따른 휴지의 실현 양상 연구)

  • Yoo, Doyoung;Shin, Jiyoung
    • Phonetics and Speech Sciences
    • /
    • v.11 no.2
    • /
    • pp.33-44
    • /
    • 2019
  • This paper dealt with how realization of pauses in utterance is affected by speech style, gender, and generation. For this purpose, we analyzed the frequency and duration of pauses. Pauses were categorized into four types: pause with breath, pause with no breath, utterance medial pause, and utterance final pause. Forty-eight subjects living in Seoul were chosen from the Korean Standard Speech Database. All subjects engaged in reading and spontaneous speech, through which we could also compare the realization between the two speech styles. The results showed that utterance final pauses had longer durations than utterance medial pauses. It means that utterance final pause has a function that signals the end of an utterance to the audience. For difference between tasks, spontaneous speech had longer and more frequent pauses because of cognitive reasons. With regard to gender variables, women produced shorter and less frequent pauses. For male speakers, the duration of pauses with breath was significantly longer. Finally, for generation variable, older speakers produced more frequent pauses. In addition, the results showed several interaction effects. Male speakers produced longer pauses, but this gender effect was more prominent at the utterance final position.

The Classification System and Information Service for Establishing a National Collaborative R&D Strategy in Infectious Diseases: Focusing on the Classification Model for Overseas Coronavirus R&D Projects (국가 감염병 공동R&D전략 수립을 위한 분류체계 및 정보서비스에 대한 연구: 해외 코로나바이러스 R&D과제의 분류모델을 중심으로)

  • Lee, Doyeon;Lee, Jae-Seong;Jun, Seung-pyo;Kim, Keun-Hwan
    • Journal of Intelligence and Information Systems
    • /
    • v.26 no.3
    • /
    • pp.127-147
    • /
    • 2020
  • The world is suffering from numerous human and economic losses due to the novel coronavirus infection (COVID-19). The Korean government established a strategy to overcome the national infectious disease crisis through research and development. It is difficult to find distinctive features and changes in a specific R&D field when using the existing technical classification or science and technology standard classification. Recently, a few studies have been conducted to establish a classification system to provide information about the investment research areas of infectious diseases in Korea through a comparative analysis of Korea government-funded research projects. However, these studies did not provide the necessary information for establishing cooperative research strategies among countries in the infectious diseases, which is required as an execution plan to achieve the goals of national health security and fostering new growth industries. Therefore, it is inevitable to study information services based on the classification system and classification model for establishing a national collaborative R&D strategy. Seven classification - Diagnosis_biomarker, Drug_discovery, Epidemiology, Evaluation_validation, Mechanism_signaling pathway, Prediction, and Vaccine_therapeutic antibody - systems were derived through reviewing infectious diseases-related national-funded research projects of South Korea. A classification system model was trained by combining Scopus data with a bidirectional RNN model. The classification performance of the final model secured robustness with an accuracy of over 90%. In order to conduct the empirical study, an infectious disease classification system was applied to the coronavirus-related research and development projects of major countries such as the STAR Metrics (National Institutes of Health) and NSF (National Science Foundation) of the United States(US), the CORDIS (Community Research & Development Information Service)of the European Union(EU), and the KAKEN (Database of Grants-in-Aid for Scientific Research) of Japan. It can be seen that the research and development trends of infectious diseases (coronavirus) in major countries are mostly concentrated in the prediction that deals with predicting success for clinical trials at the new drug development stage or predicting toxicity that causes side effects. The intriguing result is that for all of these nations, the portion of national investment in the vaccine_therapeutic antibody, which is recognized as an area of research and development aimed at the development of vaccines and treatments, was also very small (5.1%). It indirectly explained the reason of the poor development of vaccines and treatments. Based on the result of examining the investment status of coronavirus-related research projects through comparative analysis by country, it was found that the US and Japan are relatively evenly investing in all infectious diseases-related research areas, while Europe has relatively large investments in specific research areas such as diagnosis_biomarker. Moreover, the information on major coronavirus-related research organizations in major countries was provided by the classification system, thereby allowing establishing an international collaborative R&D projects.

Comparison of vowel lengths of articles and monosyllabic nouns in Korean EFL learners' noun phrase production in relation to their English proficiency (한국인 영어학습자의 명사구 발화에서 영어 능숙도에 따른 관사와 단음절 명사 모음 길이 비교)

  • Park, Woojim;Mo, Ranm;Rhee, Seok-Chae
    • Phonetics and Speech Sciences
    • /
    • v.12 no.3
    • /
    • pp.33-40
    • /
    • 2020
  • The purpose of this research was to find out the relation between Korean learners' English proficiency and the ratio of the length of the stressed vowel in a monosyllabic noun to that of the unstressed vowel in an article of the noun phrases (e.g., "a cup", "the bus", etcs.). Generally, the vowels in monosyllabic content words are phonetically more prominent than the ones in monosyllabic function words as the former have phrasal stress, making the vowels in content words longer in length, higher in pitch, and louder in amplitude. This study, based on the speech samples from Korean-Spoken English Corpus (K-SEC) and Rated Korean-Spoken English Corpus (Rated K-SEC), examined 879 English noun phrases, which are composed of an article and a monosyllabic noun, from sentences which are rated on 4 levels of proficiency. The lengths of the vowels in these 879 target NPs were measured and the ratio of the vowel lengths in nouns to those in articles was calculated. It turned out that the higher the proficiency level, the greater the mean ratio of the vowels in nouns to the vowels in articles, confirming the research's hypothesis. This research thus concluded that for the Korean English learners, the higher the English proficiency level, the better they could produce the stressed and unstressed vowels with more conspicuous length differences between them.