• Title/Summary/Keyword: Speaker Adaptation

Search Result 122, Processing Time 0.027 seconds

A Speaker Adaptation of Korean Speech Using MLLR (MLLR을 이용한 한국어 음성의 화자 적응)

  • Kim, Tae-Hyeong;Lee, Keon-Ung;Lee, Sang-Ho;Hong, Jae-Keun
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2000.10a
    • /
    • pp.251-254
    • /
    • 2000
  • 화자 독립 인식은 훈련 화자와 시험 화자의 차이로 인해 화자 종속의 경우보다 인식률이 떨어진다. 따라서, 인식률을 향상시키기 위해 화자 독립 모델을 화자에 적응시킬 필요가 있다. 본 논문에서는 효과적인 적응 방법인 MLLR(Maximum Likelihood Linear Regression) 적응 방법을 한국어 음성에 적용하여 적응 성능을 향상시켰고, 온라인 상에서 적용 가능하도록 증가 적응 방법을 이용하였다. PBW 445 음성 데이타베이스에 대한 실험 결과, 400개의 적응 데이터를 사용하였을 때, 제안한 방법이 기존의 화자 독립 시스템보다 7.02% 향상된 성능을 보였다.

  • PDF

Online Adaptation of Continuous Density Hidden Markov Models Based on Speaker Space Model Evolution (화자공간모델 진화에 근거한 연속밀도 은닉 마코프모델의 온라인 적응)

  • Kim Dong Kook;Kim Young Joon;Kim Hyun Woo;Kim Nam Soo
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • spring
    • /
    • pp.69-72
    • /
    • 2002
  • 본 논문에서 화자공간모델 evolution에 기반한 continuous density hidden Markov model (CDHMM)의 online 적응에 대한 새로운 기법을 제안한다. 학습화자의 a priori knowledge을 나타내는 화자공간모델은 factor analysis (FA) 또는 probabilistic principal component analysis (PPCA)와 같은 은닉변수모델(latent variable model)에 의해 효과적으로 나타내어진다. 은닉 변수모델은 화자공간모델뿐아니라 CDHMM 파라메터의 ajoint prior분포를 표시함으로, maximum a posteriori(MAP)적응기법에 직접 적용되어진다. 화자공간모델의 hyperparameters와 CDHMM파라메터를 동시에 순차적으로 적응하기 위해 quasi-Bayes (QB)추정 기술에 기반한 online 적응기법을 제안한다. 연속숫자음 인식과 관련된 화자적응 실험을 통해 제안된 기법은 적은 적응데이터에서 좋은 성능을 나타내며, 데이터가 증가함에 따라 성능이 지속적으로 증가함을 보여준다.

  • PDF

A study on recognition improvement of velopharyngeal insufficiency patient's speech using various types of deep neural network (심층신경망 구조에 따른 구개인두부전증 환자 음성 인식 향상 연구)

  • Kim, Min-seok;Jung, Jae-hee;Jung, Bo-kyung;Yoon, Ki-mu;Bae, Ara;Kim, Wooil
    • The Journal of the Acoustical Society of Korea
    • /
    • v.38 no.6
    • /
    • pp.703-709
    • /
    • 2019
  • This paper proposes speech recognition systems employing Convolutional Neural Network (CNN) and Long Short Term Memory (LSTM) structures combined with Hidden Markov Moldel (HMM) to effectively recognize the speech of VeloPharyngeal Insufficiency (VPI) patients, and compares the recognition performance of the systems to the Gaussian Mixture Model (GMM-HMM) and fully-connected Deep Neural Network (DNNHMM) based speech recognition systems. In this paper, the initial model is trained using normal speakers' speech and simulated VPI speech is used for generating a prior model for speaker adaptation. For VPI speaker adaptation, selected layers are trained in the CNN-HMM based model, and dropout regulatory technique is applied in the LSTM-HMM based model, showing 3.68 % improvement in recognition accuracy. The experimental results demonstrate that the proposed LSTM-HMM-based speech recognition system is effective for VPI speech with small-sized speech data, compared to conventional GMM-HMM and fully-connected DNN-HMM system.

A Focus Account for Contrastive Reduplication: Prototypicality and Contrastivity

  • Lee, Bin-Na;Lee, Chung-Min
    • Proceedings of the Korean Society for Language and Information Conference
    • /
    • 2007.11a
    • /
    • pp.259-267
    • /
    • 2007
  • This paper sets forth the phenomenon of Contrastive Reduplication (CR) in English relevant to the notion of contrastive focus (CF). CF differs from other reduplicative patterns in that rather than the general intensive function, denotation of a more prototypical and default meaning of a lexical item appears from the reduplicated form resulting as a semantic contrast with the meaning of the non-reduplicated word. Thus, CR is in concordance with CF under the concept of contrastivity. However, much of the previous works on CF associated contrastivity with a manufacture of a set of alternatives taking a semantic approach. We claim that a recent discourse-pragmatic account takes advantage of explaining the vague contrast in informativeness of CR. Zimmermann's (2006) Contrastive Focus Hypothesis characterizes contrastivity in the sense of speaker's assumptions about the hearer's expectation of the focused element. This approach makes possible adaptation to CR and recovers the possible subsets of meaning of a reduplicated form in a more refined way showing contrastivity in informativeness. Additionally, CR in other languages along with similar set-limiting phenomenon in various languages will be introduced in general.

  • PDF

Speaker Adaptation in VQ and HMM Based Speech Recognition (VQ와 HMM을 이용한 음성인식에서 화자적응에 관한 연구)

  • 이대룡
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • 1991.06a
    • /
    • pp.54-57
    • /
    • 1991
  • 본 논무에서는 HMM과 VQ를 이용한 고립단어에 대한 화자종속 및 화자독립 음성인식시스템을 만들고 여기에 화자적응을 하는 방법에 대한 연구를 했다. 화자적응방법에는 크게 VQ코드북을 적응시키는 방법과 HMM패러미터블 적응시키는 방법이 있다. 코드북적응을 하는 방법으로서 기존코드북에 대해 새로운화자의 적응음성을 양자화한 뒤 각 코드벡터에 해당하는 적응음성의 평균을 구해서 새로운 화자의 코드북을 구해주는 방법과 기준코드북에 대해 새로운화자의 적응음성을 양자화할 때 HMM의 각 상태에서 각각의 코드벡터를 발생할 확률을 거리오차의 계산에서 고려해 비록 거리오차는 크지만 그 코드벡터를 발생할 확률이 매우 높으면 적응음성이 그 코드벡터에 index되게해서 각 코드벡터에 해당하는 모든 적응음성데이타의 평균을 새로운 코드북으로 하는 두가지 알고리즘을 제안한다. 이렇게 함으로써 기존의 기준코드북을 초기 코드북으로해서 LBG알고리즘을 사용해서 적응음성데이타에 대한 새로운 코드북을 만드는 방법에 비해 5-10배의 계산시간을 감소하게 된다. 이 새로운 코드북으로 적응음성데이타를 다시 index해서 이 index된 음성렬로 HMM패러미터를 적응했다. 제안된 알고리즘이 코드북적응을 하는 경우에 기존의 적응방법에 비해 5-10배의 계산 시간을 단축하면서 인식률에서는 더 나은결과를 얻었다. 또 같은 적응방법에 대해서 화자종속모델 보다는 화자독립모델에 대해서 화자적응하는 것이 더 나은 인식결과를 보여주었다.

  • PDF

Perception of military officers towards the military adaptation of adults who stutter and the associated factors (말더듬 성인의 군대 적응 정도에 대한 군지휘관의 인식 양상 및 관련 요인 분석)

  • Hye-rin Park;Jin Park
    • Phonetics and Speech Sciences
    • /
    • v.15 no.1
    • /
    • pp.55-64
    • /
    • 2023
  • This study investigated the factors influencing the perceptions that military officers can harbor regarding persons who stutter in terms of how well they can adapt to the army. In total, 89 participants were randomly assigned to each of the three different conditions ("fluent speech"=23, "mildly stuttered speech"=34, and "severely stuttered speech"=32). Subsequently, the participants were asked to listen and rate each sample in terms of "the speaker's communicative functioning (i.e., speech fluency, intelligibility, naturalness, speech rate), personal traits (i.e., likeability, anxiety level, intellectual level, and sociability), and the perceived degree of the adaptability to the army." The results showed that significant differences were found between "fluent speech" and "severely stuttered speech" in the perceived communicative functionings and the perceived adaptability to the army. Moreover, there were significant differences in the same variables between "mildly stuttered speech" and "severely stuttered speech." However, there were no significant differences between "mildly stuttered speech" and "fluent speech." Following the conducting of the Pearson correlation test, strong correlations were also found between the perceived communicative functionings, in particular "speech fluency," and the perceived adaptability to the army. Those results can be employed to argue that the communicative functionings can serve as factors which influence the perceptions of persons who stutter in terms of how well they can adapt to the army. Further discussion has taken place regarding the relationship between the perceived communicative functionings and the perceived adaptability to the army.

Robust Speech Recognition Algorithm of Voice Activated Powered Wheelchair for Severely Disabled Person (중증 장애우용 음성구동 휠체어를 위한 강인한 음성인식 알고리즘)

  • Suk, Soo-Young;Chung, Hyun-Yeol
    • The Journal of the Acoustical Society of Korea
    • /
    • v.26 no.6
    • /
    • pp.250-258
    • /
    • 2007
  • Current speech recognition technology s achieved high performance with the development of hardware devices, however it is insufficient for some applications where high reliability is required, such as voice control of powered wheelchairs for disabled persons. For the system which aims to operate powered wheelchairs safely by voice in real environment, we need to consider that non-voice commands such as user s coughing, breathing, and spark-like mechanical noise should be rejected and the wheelchair system need to recognize the speech commands affected by disability, which contains specific pronunciation speed and frequency. In this paper, we propose non-voice rejection method to perform voice/non-voice classification using both YIN based fundamental frequency(F0) extraction and reliability in preprocessing. We adopted a multi-template dictionary and acoustic modeling based speaker adaptation to cope with the pronunciation variation of inarticulately uttered speech. From the recognition tests conducted with the data collected in real environment, proposed YIN based fundamental extraction showed recall-precision rate of 95.1% better than that of 62% by cepstrum based method. Recognition test by a new system applied with multi-template dictionary and MAP adaptation also showed much higher accuracy of 99.5% than that of 78.6% by baseline system.

Implementation of a storage device the noise elimination negative input that using adaptive filter. (적응형 필터를 이용한 잡음제거 음성입력 및 저장장치의 구현)

  • Ji, Yoo-Kang;Moon, Dae-Wong;Kim, Sa-Wung;Park, Soo-Bong
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2008.05a
    • /
    • pp.147-150
    • /
    • 2008
  • The explanation of Tourism Guide of present whole country main tourist resort helps to understand the tourist resort. However, activity space of Tourism Guide is not state that can be understood all by upbringing by a natural voice. Because action of Tourism Guide is much in case of most sightseeing explanation to use microphone and speaker etc., as sticking that attach and uses to clothing and so on uses, there are much vexatious. Treatise that see hereupon makes use of establishment style fixing microphone and embody inputted obscene sounds by On-board system inflecting MCU (ATmega128), MSM7731-02 Oki-Dual Codec to minimize noise using ecad filter, and embodied a control program by serial communication method with filter codec. The resultant audible direction the maximum 59ms, the line echo maximum 27ms, the echo decrease maximum 35dB, it embodied the system which removes the adaptation elder brother noise of the back.

  • PDF

Cognitive abilities and speakers' adaptation of a new acoustic form: A case of a /o/-raising in Seoul Korean

  • Kong, Eun Jong;Kang, Jieun
    • Phonetics and Speech Sciences
    • /
    • v.10 no.3
    • /
    • pp.1-8
    • /
    • 2018
  • The vowel /o/ in Seoul Korean has been undergoing a sound change by altering the acoustic weighting of F2 and F1. Studies documented that this on-going change redefined the nature of a /o/-/u/ contrast as F2 differences rather than as F1 differences. The current study examined two cognitive factors namely executive function capacity (EF) and autistic traits, in terms of their roles in explaining who in speech community would adapt new acoustic forms of the target vowels, and who would retain the old forms. The participants, 55 college students speaking Seoul Korean, produced /o/ and /u/ vowels in isolated words; and completed three EF tasks (Digit N-Back, Stroop, and Trail-Making Task), and an Autism screening questionnaire. The relationships between speakers' cognitive task scores and their utilizations of F1 and F2 were analyzed using a series of correlation tests. Results yielded a meaningful relationship in participants' EF scores interacting with gender. Among the females, speakers with higher EF scores were better at retaining F1, which is a less informative cue for females since they utilized F2 more than they did F1 in realizing /o/ and /u/. In contrast, better EF control among male speakers was associated with more use of the new cue (F2) where males still utilized F1 as much as F2 in the production of /o/ and /u/ vowels. Taken together, individual differences in acoustic realization can be explained by individuals' cognitive abilities, and their progress in the sound change further predicts that cognitive ability influences the utilization of acoustic information which is non-primary to the speaker.

Exploring the Study Experiences of Southeast Asian Students at a Korean University in Seoul (서울 A대학 동남아시아 유학생의 학업 경험에 대한 탐색적 연구)

  • KIM, Jeehun
    • The Southeast Asian review
    • /
    • v.23 no.3
    • /
    • pp.135-179
    • /
    • 2013
  • This study explores the study experiences of Southeast Asian students at a reputable Korean private university in Seoul. In particular, this study focuses on difficulties and coping strategies of both non-native speaker of English and native-speakers of English who are working for their undergraduate or postgraduate degrees. Interviews of fourteen students from five Southeast Asian countries were collected and analyzed by NVivo 9. Thematic analysis result shows that many students, particularly non-native speakers of English, had much more difficulties than their counterparts, in contemporary Korean university context, where internationalization indices-driven strategies including expanding courses conducted in English language. Also, this study observes and documents contrasting patterns of different degree of difficulties experienced by students, depending on their degree levels and majors. Undergraduate students in science and engineering majors had the greatest degree of difficulties among all. In contrast, their graduate counterparts seem to have less difficulties. This might be related to the fact that graduate students in science and engineering majors are mostly working with their peers in their own labs, which provides institutional support. Coping strategies of students show that international students, facing unfavorable or unfriendly treatments by their Korean peers, developed innovative strategies, including using the internet technology to catch up with the classes that they could not fully understand. As a whole, adaptation process of international students do not seem to be passive or one-way. This study also provides policy implications for international students, particularly, who can be categorized as linguistic and ethnic minorities.