Search | Korea Science

Microphone Array Based Speech Enhancement Using Independent Vector Analysis (마이크로폰 배열에서 독립벡터분석 기법을 이용한 잡음음성의 음질 개선)

Wang, Xingyang;Quan, Xingri;Bae, Keunsung
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.87-92
- /
- 2012
Speech enhancement aims to improve speech quality by removing background noise from noisy speech. Independent vector analysis is a type of frequency-domain independent component analysis method that is known to be free from the frequency bin permutation problem in the process of blind source separation from multi-channel inputs. This paper proposed a new method of microphone array based speech enhancement that combines independent vector analysis and beamforming techniques. Independent vector analysis is used to separate speech and noise components from multi-channel noisy speech, and delay-sum beamforming is used to determine the enhanced speech among the separated signals. To verify the effectiveness of the proposed method, experiments for computer simulated multi-channel noisy speech with various signal-to-noise ratios were carried out, and both PESQ and output signal-to-noise ratio were obtained as objective speech quality measures. Experimental results have shown that the proposed method is superior to the conventional microphone array based noise removal approach like GSC beamforming in the speech enhancement.
https://doi.org/10.13064/KSSS.2012.4.4.087 인용 PDF

Prosodic Disambiguation of Low versus High Syntactic Attachment across Lexical Biases in English

Jeon, Yoon-Shil;Yoon, Kyu-Chul
- Phonetics and Speech Sciences
- /
- v.4 no.1
- /
- pp.55-65
- /
- 2012
In this study, the prosodic disambiguation of the syntactic attachment differences was investigated in relation to the effect of lexical bias. Speech materials were composed of N1-conj-N2-PP phrases such as "walkers and runners with dogs." The results show that the use of durational pattern is dominant over the pitch pattern to differentiate the attachment differences. The characteristic pitch contour was the rise and fall over N1 and N2 in the high attachment. The pitch contour in the low attachment was the rise and fall over N2 and N3 although the frequency of such patterns was lower for the low attachment case. For the durational pattern, the lengthening in the N2 region plays a significant role in the disambiguation of the syntactic attachments. The interaction between the lexical bias and the syntactic attachment was not statistically significant in the duration data.
https://doi.org/10.13064/KSSS.2012.4.1.055 인용 PDF

Error Correction and Praat Script Tools for the Buckeye Corpus of Conversational Speech (벅아이 코퍼스 오류 수정과 코퍼스 활용을 위한 프랏 스크립트 툴)

Yoon, Kyu-Chul
- Phonetics and Speech Sciences
- /
- v.4 no.1
- /
- pp.29-47
- /
- 2012
The purpose of this paper is to show how to convert the label files of the Buckeye Corpus of Spontaneous Speech [1] into Praat format and to introduce some of the Praat scripts that will enable linguists to study various aspects of spoken American English present in the corpus. During the conversion process, several types of errors were identified and corrected either manually or automatically by the use of scripts. The Praat script tools that have been developed can help extract from the corpus massive amounts of phonetic measures such as the VOT of plosives, the formants of vowels, word frequency information and speech rates that span several consecutive words. The script tools can extract additional information concerning the phonetic environment of the target words or allophones.
https://doi.org/10.13064/KSSS.2012.4.1.029 인용 PDF

Perception of Korean stops with a three-way laryngeal contrast

Kong, Eun-Jong
- Phonetics and Speech Sciences
- /
- v.4 no.1
- /
- pp.13-20
- /
- 2012
A lax stop in Korean, one of the three laryngeal contrastive stops, has undergone a sound change in terms of its acoustic properties. Prior production studies described this recent lax stop as being differentiated from tense and aspirated stops primarily by fundamental frequencies (f0). And, the acoustic property of voice onset time (VOT) further separates tense stops from lax and aspirated stops. The current research explores how these two major acoustic parameters of f0 and VOT cue the three stop categories in Korean adult listeners' perception. Thirty-one native speakers of Korean participated in two experimental tasks: categorization judgment and within-category goodness ratings. Two sets of audio stimuli were prepared by synthesizing English and Korean male speakers' CV productions. The findings showed that while f0 cues listeners to lax stops as production patterns would predict, VOT were closely related to listeners' categorization and goodness ratings of lax stops. This suggests that accurate characterizations of the recent lax stop category need to be based on Korean speakers' perceptual behavior as well as production patterns.
https://doi.org/10.13064/KSSS.2012.4.1.013 인용 PDF

An Analysis of the Vowel Formants of the Young Females in the Buckeye Corpus (벅아이 코퍼스에서의 젊은 성인 여성의 모음 포먼트 분석)

Yoon, Kyuchul
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.45-52
- /
- 2012
The purpose of this paper is to measure the first two vowel formants of the ten young female speakers from the Buckeye Corpus of Conversational Speech [1] automatically and then to analyze various potential factors that may affect the formant distribution of the eight peripheral vowels of English. The factors that were analyzed included the place of articulation, the content versus function word information, the syllabic stress information, the location in a word, the location in an utterance, the speech rate of the three consecutive words, and the word frequency in the corpus. The results indicate that the overall formant patterns of the female speakers were similar to those of earlier works. The effects of the factors on the realization of the two formants were also similar to those from the male speakers with minor differences.
https://doi.org/10.13064/KSSS.2012.4.4.045 인용 PDF

An Analysis of the Vowel Formants of the Young versus Old Speakers in the Buckeye Corpus (벅아이 코퍼스에서의 연령별 모음 포먼트 분석)

Km, Ji-Eun;Yoon, Kyuchul
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.29-35
- /
- 2012
The purpose of this study was to measure the first two vowel formants of the forty male and female speakers (twenty young vs. old male speakers and twenty young vs. old female speakers) from the Buckeye Corpus of Conversational Speech and to examine the vowel formant changes across two generations (younger vs. older). The results indicated that the vowel space of the younger generation (in their thirties or less) shifted to the lower left position compared to those of the older generation (in their forties or more) in both male and female speakers. When the results were compared to those of Peterson & Barney (1952), it appears that differences can be found in the size of the vowel spaces through time.
https://doi.org/10.13064/KSSS.2012.4.4.029 인용 PDF

A comparison of the voice difference of persons with Idiopathic Parkinson's disease and a normal group in five vowels (파킨슨병 환자와 정상노인의 모음 산출 특성 비교)

Lee, In-Ae;Kim, Moon-Jeoung;Hwang, Young-Jin
- Phonetics and Speech Sciences
- /
- v.4 no.1
- /
- pp.119-124
- /
- 2012
The purpose of this study is to compare the voice differences of persons with Idiopathic Parkinson's disease and a normal group according to five vowels. Eight persons with Idiopathic Parkinson's disease and a healthy control group of 22 were selected and every voice analyzed by MDVP. The first result showed that jitter measurements between the two group showed a significant statistical difference according to all vowels. Second, the two groups' shimmer measurements showed a significant statistical difference according to nearly all vowels. Third, jitter measurements between the five vowels were more relatively closely correlated persons with Idiopathic Parkinson's disease than the normal group. Fourth, shimmer figures between the five vowels more relatively closely correlated persons with Idiopathic Parkinson's disease than the normal group.
https://doi.org/10.13064/KSSS.2012.4.1.119 인용 PDF

A Study of an Independent Evaluation of Prosody and Segmentals: with Reference to the Difference in the Foreign Accent of Korean, Chinese, and Japanese Learners of English (운율 및 분절음의 독립적 발음 평가 연구: 한국인, 중국인, 일본인 영어 학습자의 액센트 차이를 중심으로)

Park, Hansang
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.37-43
- /
- 2012
This study investigates an independent evaluation of prosody and segmentals with reference to the difference in the foreign accent of Korean, Chinese, and Japanese learners of English. For this study, a set of stimuli were made of English sentences read by male and female Korean, Chinese, and Japanese learners of English by prosody swapping technique. Two groups of American and Korean subjects evaluated the difference in the prosody and segmentals of the stimuli by pairwise difference rating. The results showed that there was no significant difference in the evaluation scores of prosody and segmentals across accents for either subject group. The results also showed that both subject groups indicated a greater score with segmentals than with prosody. The results of the present study are significant in that they are opposite to the claim of some previous studies that prosodic factors could have a greater influence on the foreign accent and intelligibility than segmentals.
https://doi.org/10.13064/KSSS.2012.4.4.037 인용 PDF

Speech Rate Variation in Synchronous Speech (동시발화에 나타나는 발화 속도 변이 분석)

Kim, Miran;Nam, Hosung
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.19-27
- /
- 2012
When two speakers read a text together, the produced speech has been shown to reduce a high degree of variability (e.g., pause duration and placement, and speech rate). This paper provides a quantitative analysis of speech rate variation exhibited in synchronous speech by examining the global and local patterns in two dialects of Mandarin Chinese (Taiwan and Shanghai). We analyzed the speech data in terms of mean speech rate and the reference of "Just Noticeable difference (JND)" within a subject and across subjects. Our findings show that speakers show lower and less variable speech rates when they read a text synchronously than when they read alone. This global pattern is observed consistently across speakers and dialects maintaining the unique local variation patterns of speech rate for each dialect. We conclude that paired speakers lower their speech rates and decrease the variability in order to ensure the synchrony of their speech.
https://doi.org/10.13064/KSSS.2012.4.4.019 인용 PDF

Harmonic Peak Picking-based MVF Estimation for Improvement of HMM-based Speech Synthesis System Using TBE Model (TBE 모델을 사용하는 HMM 기반 음성합성기 성능 향상을 위한 하모닉 선택에 기반한 MVF 예측 방법)

Park, Jihoon;Hahn, Minsoo
- Phonetics and Speech Sciences
- /
- v.4 no.4
- /
- pp.79-86
- /
- 2012
In the two-band excitation (TBE) model, maximum voiced frequency (MVF) is the most important feature of the excitation parameter because the synthetic speech quality depends on MVF. Thus, this paper proposes an enhanced MVF estimation scheme based on the peak picking method. In the proposed scheme, the local peak and the peak lobe are picked from the spectrum of a linear predictive residual signal. The normalized distance between neighboring peak lobes is calculated and utilized as a feature to estimate MVF. Experimental results of both objective and subjective tests show that the proposed scheme improves synthetic speech quality compared with that of the conventional one.
https://doi.org/10.13064/KSSS.2012.4.4.079 인용 PDF

Search Result 948, Processing Time 0.019 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)