통합 검색 | Korea Science

Digital enhancement of pronunciation assessment: Automated speech recognition and human raters

Miran Kim
- 말소리와 음성과학
- /
- 제15권2호
- /
- pp.13-20
- /
- 2023
This study explores the potential of automated speech recognition (ASR) in assessing English learners' pronunciation. We employed ASR technology, acknowledged for its impartiality and consistent results, to analyze speech audio files, including synthesized speech, both native-like English and Korean-accented English, and speech recordings from a native English speaker. Through this analysis, we establish baseline values for the word error rate (WER). These were then compared with those obtained for human raters in perception experiments that assessed the speech productions of 30 first-year college students before and after taking a pronunciation course. Our sub-group analyses revealed positive training effects for Whisper, an ASR tool, and human raters, and identified distinct human rater strategies in different assessment aspects, such as proficiency, intelligibility, accuracy, and comprehensibility, that were not observed in ASR. Despite such challenges as recognizing accented speech traits, our findings suggest that digital tools such as ASR can streamline the pronunciation assessment process. With ongoing advancements in ASR technology, its potential as not only an assessment aid but also a self-directed learning tool for pronunciation feedback merits further exploration.
https://doi.org/10.13064/KSSS.2023.15.2.013 인용 PDF

How Korean Learner's English Proficiency Level Affects English Speech Production Variations

Hong, Hye-Jin;Kim, Sun-Hee;Chung, Min-Hwa
- 말소리와 음성과학
- /
- 제3권3호
- /
- pp.115-121
- /
- 2011
This paper examines how L2 speech production varies according to learner's L2 proficiency level. L2 speech production variations are analyzed by quantitative measures at word and phone levels using Korean learners' English corpus. Word-level variations are analyzed using correctness to explain how speech realizations are different from the canonical forms, while accuracy is used for analysis at phone level to reflect phone insertions and deletions together with substitutions. The results show that speech production of learners with different L2 proficiency levels are considerably different in terms of performance and individual realizations at word and phone levels. These results confirm that speech production of non-native speakers varies according to their L2 proficiency levels, even though they share the same L1 background. Furthermore, they will contribute to improve non-native speech recognition performance of ASR-based English language educational system for Korean learners of English.
PDF

영어의 강음절(강세 음절)과 한국어 화자의 단어 분절 (Strong (stressed) syllables in English and lexical segmentation by Koreans)

김선미;남기춘
- 말소리와 음성과학
- /
- 제3권1호
- /
- pp.3-14
- /
- 2011
It has been posited that in English, native listeners use the Metrical Segmentation Strategy (MSS) for the segmentation of continuous speech. Strong syllables tend to be perceived as potential word onsets for English native speakers, which is due to the high proportion of strong syllables word-initially in the English vocabulary. This study investigates whether Koreans employ the same strategy when segmenting speech input in English. Word-spotting experiments were conducted using vowel-initial and consonant-initial bisyllabic targets embedded in nonsense trisyllables in Experiment 1 and 2, respectively. The effect of strong syllable was significant in the RT (reaction times) analysis but not in the error analysis. In both experiments, Korean listeners detected words more slowly when the word-initial syllable is strong (stressed) than when it is weak (unstressed). However, the error analysis showed that there was no effect of initial stress in Experiment 1 and in the item (F2) analysis in Experiment 2. Only the subject (F1) analysis in Experiment 2 showed that the participants made more errors when the word starts with a strong syllable. These findings suggest that Koran listeners do not use the Metrical Segmentation Strategy for segmenting English speech. They do not treat strong syllables as word beginnings, but rather have difficulties recognizing words when the word starts with a strong syllable. These results are discussed in terms of intonational properties of Korean prosodic phrases which are found to serve as lexical segmentation cues in the Korean language.
PDF

Speech recognition rates and acoustic analyses of English vowels produced by Korean students

Yang, Byunggon
- 말소리와 음성과학
- /
- 제14권2호
- /
- pp.11-17
- /
- 2022
English vowels play an important role in verbal communication. However, Korean students tend to experience difficulty pronouncing a certain set of vowels despite extensive education in English. The aim of this study is to apply speech recognition software to evaluate Korean students' pronunciation of English vowels in minimal pair words and then to examine acoustic characteristics of the pairs in order to check their pronunciation problems. Thirty female Korean college students participated in the recording. Speech recognition rates were obtained to examine which English vowels were correctly pronounced. To compare and verify the recognition results, such acoustic analyses as the first and second formant trajectories and durations were also collected using Praat. The results showed an overall recognition rate of 54.7%. Some students incorrectly switched the tense and lax counterparts and produced the same vowel sounds for qualitatively different English vowels. From the acoustic analyses of the vowel formant trajectories, some of these vowel pairs were almost overlapped or exhibited slight acoustic differences at the majority of the measurement points. On the other hand, statistical analyses on the first formant trajectories of the three vowel pairs revealed significant differences throughout the measurement points, a finding that requires further investigation. Durational comparisons revealed a consistent pattern among the vowel pairs. The author concludes that speech recognition and analysis software can be useful to diagnose pronunciation problems of English-language learners.
https://doi.org/10.13064/KSSS.2022.14.2.011 인용 PDF KSCI

Acoustic analysis of English lexical stress produced by Korean, Japanese and Taiwanese-Chinese speakers

Jung, Ye-Jee;Rhee, Seok-Chae
- 말소리와 음성과학
- /
- 제10권1호
- /
- pp.15-22
- /
- 2018
Stressed vowels in English are usually produced using longer duration, higher pitch, and greater intensity than unstressed vowels. However, many English as a foreign language (EFL) learners have difficulty producing English lexical stress because their mother tongues do not have such features. In order to investigate if certain non-native English speakers (Korean, Japanese, and Taiwanese-Chinese native speakers) are able to produce English lexical stress in a native-like manner, speech samples were extracted from the L2 learners' corpus known as AESOP (the Asian English Speech cOrpus Project). Sixteen disyllabic words were analyzed in terms of the ratio of duration, pitch, and intensity. The results demonstrate that non-native English speakers are able to produce English stress in a similar way to native English speakers, and all speakers (both native and non-native) show a tendency to use duration as the strongest cue in producing stress. The results also show that the duration ratio of native English speakers was significantly higher than that of non-native speakers, indicating that native speakers produce a bigger difference in duration between stressed and unstressed vowels.
https://doi.org/10.13064/KSSS.2018.10.1.015 인용 PDF KSCI

영어의 억양 유형화를 이용한 발화 속도와 남녀 화자에 따른 음향 분석 (An acoustical analysis of speech of different speaking rates and genders using intonation curve stylization of English)

이서배
- 말소리와 음성과학
- /
- 제6권4호
- /
- pp.79-90
- /
- 2014
An intonation curve stylization was used for an acoustical analysis of English speech. For the analysis, acoustical feature values were extracted from 1,848 utterances produced with normal and fast speech rate by 28 (12 women and 16 men) native speakers of English. Men are found to speak faster than women at normal speech rate but no difference is found between genders at fast speech rate. Analysis of pitch point features has it that fast speech has greater Pt (pitch point movement time), Pr (pitch point pitch range), and Pd (pitch point distance) but smaller Ps (pitch point slope) than normal speech. Men show greater Pt, Pr, and Pd than women. Analysis of sentence level features reveals that fast speech has smaller Sr (sentence level pitch range), Sd (sentence duration), and Max (maximum pitch) but greater Ss (sentence slope) than normal speech. Women show greater Sr, Ss, Sp (pitch difference between the first pitch point and the last), Sd, MaxNr (normalized Max), and MinNr (normalized Min) than men. As speech rate increases, women speak with greater Ss and Sr than men.
https://doi.org/10.13064/KSSS.2014.6.4.079 인용 PDF KSCI

한국 중학생의 영어 읽기 발화에서 문장유형에 따른 유창성 등급과 초분절 요소의 관계 (The relationship between fluency levels and suprasegmentals according to the sentence types in the English read speech by Korean middle school English learners)

김화영
- 말소리와 음성과학
- /
- 제14권3호
- /
- pp.51-66
- /
- 2022
본 연구의 목적은 한국인 영어 학습자가 영어문장을 읽을 때 어떠한 초분절 요소가 영어 원어민 화자에 가깝게 구현되는데 영향을 미치는지를 밝혀 영어 발음교육에 도움이 되고자 하는 것이다. 본 연구에서는 연구대상자를 중학생 영어학습자로 선택하고, 다양한 유형의 문장(평서문, 의문문, 명령문, 감탄문)과 음절수로 연구 자료를 구성하였다. 이들 영어 문장 발화의 분석대상으로는 초분절 요소 중 발화속도, 휴지빈도, 휴지길이, F0 범위, 리듬을 이용하였고 음성분석 결과는 평균분석, 상관분석 및 회귀분석을 실시하였다. 그 결과, 발화속도, 휴지빈도, 휴지길이, F0 범위가 유창성 등급 평가에 영향을 미친다는 결과를 얻었다. 모든 초분절 요소와 유창성 등급 간의 회귀분석에서는 유창성 등급에 영향을 미치는 초분절 요소는 발화속도와 F0 범위이다. 리듬은 유창성 등급과의 관계에서 통계적으로 유의미하지 않았다. 따라서, 영어 발음교육을 할 때 발화속도를 높이고, F0 범위를 크게 하도록 교육하는 것이 필요하다. 또한, 발화시 휴지개수와 휴지시간을 줄이도록 하는 교육이 유창성을 높이는데 도움이 된다. 문장유형을 분류하여 분석한 결과, 감탄문의 경우 다른 문장유형에 비해 발화속도가 더 빠르고, 휴지빈도는 더 적고, 휴지길이는 더 짧으며, 리듬값은 더 높았다.
https://doi.org/10.13064/KSSS.2022.14.3.051 인용 PDF KSCI

A Corpus-based Lexical Analysis of the Speech Texts: A Collocational Approach

Kim, Nahk-Bohk
- 영어어문교육
- /
- 제15권3호
- /
- pp.151-170
- /
- 2009
Recently speech texts have been increasingly used for English education because of their various advantages as language teaching and learning materials. The purpose of this paper is to analyze speech texts in a corpus-based lexical approach, and suggest some productive methods which utilize English speaking or writing as the main resource for the course, along with introducing the actual classroom adaptations. First, this study shows that a speech corpus has some unique features such as different selections of pronouns, nouns, and lexical chunks in comparison to a general corpus. Next, from a collocational perspective, the study demonstrates that the speech corpus consists of a wide variety of collocations and lexical chunks which a number of linguists describe (Lewis, 1997; McCarthy, 1990; Willis, 1990). In other words, the speech corpus suggests that speech texts not only have considerable lexical potential that could be exploited to facilitate chunk-learning, but also that learners are not very likely to unlock this potential autonomously. Based on this result, teachers can develop a learners' corpus and use it by chunking the speech text. This new approach of adapting speech samples as important materials for college students' speaking or writing ability should be implemented as shown in samplers. Finally, to foster learner's productive skills more communicatively, a few practical suggestions are made such as chunking and windowing chunks of speech and presentation, and the pedagogical implications are discussed.
PDF

영어 동시발화의 자동 억양궤적 추출을 통한 음향 분석 (An acoustical analysis of synchronous English speech using automatic intonation contour extraction)

이서배
- 말소리와 음성과학
- /
- 제7권1호
- /
- pp.97-105
- /
- 2015
This research mainly focuses on intonational characteristics of synchronous English speech. Intonation contours were extracted from 1,848 utterances produced in two different speaking modes (solo vs. synchronous) by 28 (12 women and 16 men) native speakers of English. Synchronous speech is found to be slower than solo speech. Women are found to speak slower than men. The effect size of speech rate caused by different speaking modes is greater than gender differences. However, there is no interaction between the two factors (speaking modes vs. gender differences) in terms of speech rate. Analysis of pitch point features has it that synchronous speech has smaller Pt (pitch point movement time), Pr (pitch point pitch range), Ps (pitch point slope) and Pd (pitch point distance) than solo speech. There is no interaction between the two factors (speaking modes vs. gender differences) in terms of pitch point features. Analysis of sentence level features reveals that synchronous speech has smaller Sr (sentence level pitch range), Ss (sentence slope), MaxNr (normalized maximum pitch) and MinNr (normalized minimum pitch) but greater Min (minimum pitch) and Sd (sentence duration) than solo speech. It is also shown that the higher the Mid (median pitch), the MaxNr and the MinNr in solo speaking mode, the more they are reduced in synchronous speaking mode. Max, Min and Mid show greater speaker discriminability than other features.
https://doi.org/10.13064/KSSS.2015.7.1.097 인용 PDF KSCI

벅아이 코퍼스를 이용한 영어 무성파열음의 VOT 연구 (A Study on the Voice Onset Time of English Voiceless Stops in the Buckeye Corpus)

윤규철
- 말소리와 음성과학
- /
- 제4권2호
- /
- pp.33-40
- /
- 2012
The purpose of this paper is to investigate the voice onset time (VOT) of the English voiceless stops [p, t, k] found in the Buckeye Corpus of Conversational Speech [1]. Three young female speakers were chosen for this study and their VOT values were semi-automatically extracted along with other factors. The factors used for the analysis were place of articulation, location in word, syllabic stress, content word or not, word frequency calculated from the corpus, and the speech rate expressed in syllables per second. Results showed that, for the three places of articulation of each speaker, all the factors had a statistically significant effect on the VOT values. This paper has significance in that the materials used for the analysis were from a corpus of spontaneous natural English speech.
https://doi.org/10.13064/KSSS.2012.4.2.033 인용 PDF

검색결과 162건 처리시간 0.019초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)