• Title/Summary/Keyword: fundamental frequency of speech

Search Result 205, Processing Time 0.022 seconds

A study on speech training aids for Deafs (청각장애자용 발음훈련기기 개발에 관한 연구)

  • Ahn, Sang-Pil;Lee, Jae-Hyuk;Yoon, Tae-Sung;Park, Sang-Hui
    • Proceedings of the KIEE Conference
    • /
    • 1990.07a
    • /
    • pp.47-50
    • /
    • 1990
  • Deafs cannot speak straight voice as normal people in lack of feedback of their pronunciation, therefore speech training is required. In this study, fundamental frequency, intensity, formant frequencies, vocal tract graphic and vocal tract area function, extracted from speech signal, are used as feature parameter. AR model, whose coefficients are extracted using inverse filtering. is used as speech generation model. In connect ion between vocal tract graphic and speech parameter, articulation distances and articulation distance functions in selected 15-intervals are determined by extracted vocal tract areas and formant frequencies.

  • PDF

Diagnosis and Evaluation of Humanities Therapy: The Phonetic Analysis of Speech Rates and Fundamental Frequency According to Preferred Sensation Type (인문치료의 진단 및 평가: 감각유형에 따른 말속도와 기본주파수의 실험음성학적 분석)

  • Lee, Chan-Jong;Heo, Yun-Ju
    • The Journal of the Acoustical Society of Korea
    • /
    • v.30 no.4
    • /
    • pp.231-237
    • /
    • 2011
  • The purpose of this study is to examine the correlation between the preferred sensation type and speech sounds, especially on $F_0$ and the speech rates. Data for the sensation types and speech sounds were collected from 36 undergraduate and graduate students (17 male, 19 female). Subjects were asked to read a given text (400 syllables), describe a drawing, and give answers to some questions. We measured speakers' $F_0$ and speech rates. The results show that type V (Visual) has the correlation with the speech rates when type D (Digital) was ruled out, and type A (Auditory) has the correlation with the speech rates when type D was included. Furthermore, the analysis of the mean values of V, A, K (Visual, Auditory, Kinethetic) indicates that type V is characterized with faster speech rates and higher $F_0$ in all parts except for interview and the same is true for that of V, A, K, D (Visual, Auditory, Kinethetic, Digital) in all parts. In conclusion, this study proved that the preferred sensation type has the correlation with $F_0$ and speech rates. Based on the results of this study, $F_0$ and speech rates can be used to analyze the sensation types for individualized education as well as consultation. In addition, this study has great significance in that it lays a foundation for the study on the correlation between a preferred sensation type and speech sounds.

Pitch Estimation Method in an Integrated Time and Frequency Domain by Applying Linear Interpolation (선형 보간법을 이용한 시간과 주파수 조합영역에서의 피치 추정 방법)

  • Kim, Ki-Chul;Park, Sung-Joo;Lee, Seok-Pil;Kim, Moo-Young
    • Journal of the Institute of Electronics Engineers of Korea SP
    • /
    • v.47 no.5
    • /
    • pp.100-108
    • /
    • 2010
  • An autocorrelation method is used in pitch estimation. Autocorrelation values in time and frequency domains, which have different characteristics, correspond to the pitch period and fundamental frequency, respectively. We utilize an integrated autocorrelation method in time and frequency domains. It can remove the errors of pitch doubling and having. In the time and frequency domains, pitch period and fundamental frequency have reciprocal relation to each other. Especially, fundamental frequency estimation ends up as an error because of the resolution of FFT. To reduce these artifacts, interpolation methods are applied in the integrated autocorrelation domain, which decreases pitch errors. Moreover, only for the pitch candidates found in a time domain, the corresponding frequency-domain autocorrelation values are calculated with reduced computational complexity. Using linear interpolation, we can decrease the required number of FFT coefficients by 8 times. Thus, compared to the conventional methods, computational complexity can be reduced by 9.5 times.

A New Stylization Method using Least-Square Error Minimization on Segmental Pitch Contour (최소 자승오차 방식을 이용한 세그먼트 피치패턴의 정형화)

  • 이정철
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • 1994.06c
    • /
    • pp.107-110
    • /
    • 1994
  • In this paper, we describe the features of the fundamental frequency contour of Korean read speech, and propose a new stylization method to characterize the Fø pattern of segments. Our algorithm consists of three stylization processes : the segment level, the syllable level, and the sord level. For stylization of Fø contour in the segment level , we applied least square error minimization method to determine Fø values at initial, medial, and final position in a segment. In the syllable level, we determine the stylized Fø pattern of a syllable using the mean Fø value of each word and style information for each word, syllable and segment, we reconstruct Fø contour of sentences. The simulation results show that the error is less than 10% of the actual Fø contour for each sentence. In perception test, there is little difference between the synthesized speech with the original difference between the synthesized speech with the original Fø contour and the synthesized speech with the stylized Fø contour.

  • PDF

Acoustic Variation in infant crying (아기 울음의 음향학적 특성)

  • Choi, Yoon-Mi;Kim, Sun-Jun;Joo, Chan-Uhng;Kim, Hyun-Gi
    • Proceedings of the KSPS conference
    • /
    • 2007.05a
    • /
    • pp.146-148
    • /
    • 2007
  • Studies of cry characteristics in the newborn infant were aimed to determine if cry analysis could be succesful in the early detection of the infant at risk for developmental difficulties. Crying presupposes functioning of the respiratory, laryngeal and supralaryngeal muscles. The nervous system controls the capacity, stability, and co-ordination of the movements in these muscles. Hence, the cry provides information about how the Nervous System is functioning. 3 patients(down syndrome, cornelia de lange syndrome, Patent ductus arteriosus) were assessed through a Computerized Speech Lab (CSL). Tests had been chosen to assess Fundamental frequency(mean, maximum, minimum values), Melody contour, NHR, Energy. We compared the data from patients and healthy volunteer. Variations in cry characteristics were documented in a number of medical abnormalities.

  • PDF

The Comparison of Pitch Production Between Children with Cochlear Implants and Normal Hearing Children

  • Yoo, Hyun-Soo;Ko, Do-Heung
    • Speech Sciences
    • /
    • v.15 no.1
    • /
    • pp.87-98
    • /
    • 2008
  • This study compares the pitch production of children using cochlear implants (CI) with that of children with normal hearing. Twenty subjects from six to eight years old participated in the study. Three kinds of sentences were read and analyzed using Visi-Pitch $\blacktriangleright$(KAY Elemetrics, Model 3300). There were no considerable differences between the two groups regarding pitch, mean fundamental frequency (F0) and pitch range. In the cases of the slope value of F0 and duration, however, there were significant differences. Thus, it is concluded that duration and pitch control can be crucial factors in determining the intonation treatment of the children with cochlear implants.

  • PDF

The Acoustic Study on the Voices of Chines Normal Adults (중국 성인의 음성에 관한 기본 음성 측정치 연구)

  • Kim, Ji-Chae;Jeong, Ok-Ran
    • Proceedings of the KSPS conference
    • /
    • 2007.05a
    • /
    • pp.163-166
    • /
    • 2007
  • Our present study was performed to investigate acoustically the Chines normal adults' voices. 60 Chines normal adults (30 males and 30 females) of the age of 20 to 39 years oridyced systained vowel /a/ and, by analyzing them acoustically with Dr. Speech, we could get the fundamental frequency (Fo), jitter, shimmer, NNE. As results, on the average, male voices showed 1I8.1Hz in Fo, 0.186% in jitter, 1.12% in shimmer, and -13.7dB in NNE. And, female voices showed 252.4Hz in Fo, 0.186% in jitter, 0.81% in shimmer, and -1I.3dB in NNE. Every parameter except Fo showed no significant difference between male and female voices.

  • PDF

The Production and Perception of the Korean Stops by English Learners (영어권 화자의 국어 폐쇄음 발화와 지각)

  • Kim, Kee-Ho;Park, Yoon-Jin;Chun, Yun-Sil
    • Speech Sciences
    • /
    • v.13 no.4
    • /
    • pp.51-67
    • /
    • 2006
  • This study examined the acoustic properties of initial stops in Korean, produced by Korean native speakers and English Korean learners. The productions of Korean native speakers were compared with those of beginners and advanced learners of Korean. Fundamental frequency(F0) and Voice Onset Time(VOT) were measured in condition of one or two syllable words, containing word-initial lenis, fortis, and aspirated stops. English Korean Learners showed that they produced stops with relatively shorter VOT and lower F0, compared with those of Korean native speakers. In case of the manner of articulation, English Korean learners have production difficulties in order of lenis stops, aspirated stops, and fortis stops. In regard to the place of articulation, English Korean learners showed production troubles in order of labial stops, velar stops, and alveolar stops. In the experiment of perception, it is hard for English Korean learners to distinguish stops of lenis and aspirated. Therefore, the results of production experiment were almost consistent with those of the perception experiment. Finally, according to both groups of proficiency, the results demonstrated that the advanced learners produce or perceive Korean stops easier than the beginners.

  • PDF

An Acoustic Study of English Sentence Stress and Rhythm Produced by Korean Speakers

  • Kim, Ok-Young
    • Speech Sciences
    • /
    • v.14 no.1
    • /
    • pp.121-135
    • /
    • 2007
  • The purpose of this paper is to examine how Korean speakers realize English stress and rhythm at the sentence level, and investigate what different acoustic characteristics of English sentence stress and rhythm Korean speakers have, compared with those of American English speakers. Stressed words in the sentence were analyzed in terms of duration, fundamental frequency, and intensity of the stressed vowel in the word with neutral stress and with emphatic stress, respectively. According to the results, when the words had emphatic stress, both Koreans' and Americans' F0 and intensity of the stressed vowel were higher than those with neutral stress. Korean speakers of English realized the sentence stress with shorter vowel duration and higher F0 than American English speakers when the words had emphatic stress. The analysis of the timing of the sentence with increased unstressed syllables showed that both Americans and Koreans produced the sentence with longer duration as the number of unstressed syllables increased. However, the duration of unstressed syllables between stressed syllables by Koreans was longer than that by Americans. Americans seemed to produce unstressed syllables between stressed syllables faster than Koreans for regular intervals of stressed syllables. This analysis implies that if there are more unstressed syllables between stressed syllables, Koreans might produce unstressed syllables and the whole sentence with longer duration.

  • PDF

Acoustic and Physiologic Characteristics of Newborn Infants' Communication Intent via Crying (신생아 울음의 의사소통 의도와 관련된 음향학적 특성)

  • Jang, Hyo-Ryung;Ko, Do-Heung
    • Phonetics and Speech Sciences
    • /
    • v.5 no.3
    • /
    • pp.55-60
    • /
    • 2013
  • The purpose of this study was to investigate the acoustic characteristics of crying infants according to the communication intents such as hunger and pain in terms of acoustic differences in the fundamental frequency ($F_0$), jitter, shimmer, noise-to-harmonic ratio(NHR), habitual pitch, and intensity. The subjects were 20 healthy, normal infants, less than seven days old, from the city of Seoul and were born after 38 to 42 weeks(full term) of pregnancy. The sound of crying was recorded for three minutes. The crying due to pain was induced by means of the inborn metabolism error test, whereas the crying due to hunger was verified by means of the rooting reflex by waiting for the designated eating time. The results were as follows: (1) the fundamental frequency, noise-to-harmonic ratio(NHR), and intensity of the infants' crying due to pain was higher than that by hunger, showing a significant difference between the mean values. (2) the infants' crying due to hunger and that by pain did not have a significant difference in the mean jitter and shimmer values but both of them were largely outside of the normal threshold values(jitter by 1.04% and shimmer by 3.81%). This study was significant in the sense that it showed the acoustic characteristics of infants' crying from hunger and pain were very different from each other according to the communication intents in terms of the six acoustic parameters.