• Title/Summary/Keyword: Speech function

Search Result 696, Processing Time 0.027 seconds

A Study on the Robust Pitch Period Detection Algorithm in Noisy Environments (소음환경에 강인한 피치주기 검출 알고리즘에 관한 연구)

  • Seo Hyun-Soo;Bae Sang-Bum;Kim Nam-Ho
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2006.05a
    • /
    • pp.481-484
    • /
    • 2006
  • Pitch period detection algorithms are applied to various speech signal processing fields such as speech recognition, speaker identification, speech analysis and synthesis. Furthermore, many pitch detection algorithms of time and frequency domain have been studied until now. AMDF(average magnitude difference function) ,which is one of pitch period detection algorithms, chooses a time interval from the valley point to the valley point as the pitch period. AMDF has a fast computation capacity, but in selection of valley point to detect pitch period, complexity of the algorithm is increased. In order to apply pitch period detection algorithms to the real world, they have robust prosperities against generated noise in the subway environment etc. In this paper we proposed the modified AMDF algorithm which detects the global minimum valley point as the pitch period of speech signals and used speech signals of noisy environments as test signals.

  • PDF

English Conversation System Using Artificial Intelligent of based on Virtual Reality (가상현실 기반의 인공지능 영어회화 시스템)

  • Cheon, EunYoung
    • Journal of the Korea Convergence Society
    • /
    • v.10 no.11
    • /
    • pp.55-61
    • /
    • 2019
  • In order to realize foreign language education, various existing educational media have been provided, but there are disadvantages in that the cost of the parish and the media program is high and the real-time responsiveness is poor. In this paper, we propose an artificial intelligence English conversation system based on VR and speech recognition. We used Google CardBoard VR and Google Speech API to build the system and developed artificial intelligence algorithms for providing virtual reality environment and talking. In the proposed speech recognition server system, the sentences spoken by the user can be divided into word units and compared with the data words stored in the database to provide the highest probability. Users can communicate with and respond to people in virtual reality. The function provided by the conversation is independent of the contextual conversations and themes, and the conversations with the AI assistant are implemented in real time so that the user system can be checked in real time. It is expected to contribute to the expansion of virtual education contents service related to the Fourth Industrial Revolution through the system combining the virtual reality and the voice recognition function proposed in this paper.

Noise-Biased Compensation of Minimum Statistics Method using a Nonlinear Function and A Priori Speech Absence Probability for Speech Enhancement (음질향상을 위해 비선형 함수와 사전 음성부재확률을 이용한 최소통계법의 잡음전력편의 보상방법)

  • Lee, Soo-Jeong;Lee, Gang-Seong;Kim, Sun-Hyob
    • The Journal of the Acoustical Society of Korea
    • /
    • v.28 no.1
    • /
    • pp.77-83
    • /
    • 2009
  • This paper proposes a new noise-biased compensation of minimum statistics(MS) method using a nonlinear function and a priori speech absence probability(SAP) for speech enhancement in non-stationary noisy environments. The minimum statistics(MS) method is well known technique for noise power estimation in non-stationary noisy environments. It tends to bias the noise estimate below that of true noise level. The proposed method is combined with an adaptive parameter based on a sigmoid function and a priori speech absence probability (SAP) for biased compensation. Specifically. we apply the adaptive parameter according to the a posteriori SNR. In addition, when the a priori SAP equals unity, the adaptive biased compensation factor separately increases ${\delta}_{max}$ each frequency bin, and vice versa. We evaluate the estimation of noise power capability in highly non-stationary and various noise environments, the improvement in the segmental signal-to-noise ratio (SNR), and the Itakura-Saito Distortion Measure (ISDM) integrated into a spectral subtraction (SS). The results shows that our proposed method is superior to the conventional MS approach.

Analysis of acoustical characteristic changes in voice after drinking and singing (음주 및 가창 후 음성의 음향학적 특성 변화 분석)

  • Hwang, Bo-Myung;Noh, Dong-Woo;Paik, Eun-A;Jeong, Ok-Ran
    • Speech Sciences
    • /
    • v.8 no.2
    • /
    • pp.39-48
    • /
    • 2001
  • The purpose of this study was to examine changes in acoustic characteristics after drinking alcoholic beverages and singing in order to establish guidelines for vocal hygiene of both singers and non-singers. 21 university students (10 males and 11 females) vocalized /a/ before drinking, after drinking and after singing. Changes in vocal range and acoustic characteristics were analyzed by Dr. Speech 4.0 (Tigers Electronics). No significant difference was observed in vocal range following drinking. However, there was statistically significant changes in vocal range after singing. We may infer that appropriate amount of singing functioning as vocal warm-up, rather than drinking alone, resulted in improvement in their abilities to lengthen vocal folds. This is directly related to the ability to produce high-pitched sounds. Changes in jitter in female voices after singing was the only acoustic factor that was significant. Changes in Shimmer and NNE was not significant either after drinking nor singing. Subjects who were judged to perform better in singing were marked by minimum acoustic changes, which may due to their well-trained vocal fold function. The results of this study may address the necessity for vocal function exercises for the patients with neurogenic voice disorders including dysarthria. The need for more extensive research with a larger number of subjects including professional voice users is also addressed.

  • PDF

A Study on Vowel Formant Variation by Vocal Tract Modification (성도 변형에 따른 모음 포먼트의 변화 고찰)

  • Yang, Byung-Gon
    • Speech Sciences
    • /
    • v.3
    • /
    • pp.83-92
    • /
    • 1998
  • Vowels are classified by vocal tract shapes. These shapes form constriction points along the tract, which have an influence on such vocal tract resonance as $F_l,\;F_2,\;F_3$, and so on. This study reviews the perturbation theory of the tract and determines the corresponding formant frequencies from modified vocal tracts using vocal tract area function. Then, formant variation is observed from the theory. Finally, each set of $F_l,\;F_2,\;and\;F_3$ frequency is input to a speech synthesis software to make a vowel sound. Auditory impression of each sound without any modification of its vocal tract shape is almost the same as the corresponding phonetic symbol. Formant frequencies of $F_l,\;F_2,\;F_3$ vary according to the perturbation theory. Generally, constriction along the node causes formant values to decrease while constriction along the anti-node cause it to increase. Vocal tracts modified by more than $3\;cm^2$ change vowel qualities of /a/ and /i/ into those of f /v/ and /$\varepsilon$/, respectively. This study will be helpful in simulating sounds from modified vocal tracts before any operation. Further studies are desirable to compare vocal tract shapes of various languages and their sounds together.

  • PDF

A Simple and Fast Pitch Search Algorithm Using a Modified Skipping Technique in CELP Vocoder (개선된 Skipping 기법을 이용한 CELP 보코더에서의 고속피치검색 알고리듬)

  • Lee, Joo-Hun;Bae, Myung-Jin;Kwon, Choon-Woo
    • The Journal of the Acoustical Society of Korea
    • /
    • v.14 no.2E
    • /
    • pp.33-36
    • /
    • 1995
  • Based on the Characteristics of the correlation function of speech signal, the skipping technique can reduced the computation time considerably with a little degradation of speech quality. To improve the speech quality of the skipping technique, we use the reduced form of the correlation function to check the sign of the correlation value before the match score is calculated. The experimental results show that this modified skipping technique can reduce the computation time in pitch search over 35% compared with the traditional full search method without quality degradation.

  • PDF

The Effect of Vocal Function Exercise on Voice Improvement in Patients with Vocal Nodules (성대 기능 훈련이 성대결절 환자의 음성개선에 미치는 효과)

  • Lim, Hye-Jin;Kim, Jeong-Kyu;Kwon, Do-Ha;Park, Jun-Young
    • Phonetics and Speech Sciences
    • /
    • v.1 no.2
    • /
    • pp.37-42
    • /
    • 2009
  • The purpose of the present study was to determine the effect of the management program known as vocal function exercise (VFE) on voice quality. Typical VFE was modified and applied to patients with vocal nodules by controlling intensity of voice and relieving the vocal fold to solve hyperfunctional problems in VFE. Eight female subjects aged between 28 and 54 who had been diagnosed with vocal nodules took part in the study. The patients performed VFEs once a week for eight weeks. Vocal function exercises consist of voice hygiene, respiratory training, phonation training, and glide training. The subjects' voices were analyzed pre and post therapy on the aspects of acoustics, maximum phonation time (MPT), GRBAS, and voice handicap index (VHI). As a result, it was found that fundamental frequency ($F_o$) was significant increased, shimmer decreased remarkably and that noise to harmonic ratio (NHR) lowered obviously in the acoustic parameter. In addition, MPT was increased significantly. The scale of GRBAS indicated significant improvement in grade, roughness, and strained voice. VHI indicated significant improvement in an emotional part. In conclusion, VFE was effective in improving voice quality for patients with vocal nodules.

  • PDF

The Internal Structure of an Identification Function in Korean Lexical Pitch Accent in North Kyungsang Dialect

  • Kim, Jungsun
    • Phonetics and Speech Sciences
    • /
    • v.5 no.1
    • /
    • pp.91-98
    • /
    • 2013
  • This paper investigated Korean prosody as it relates to graded internal structure in an identification function. Within Korean prosody, variants regarded as dialectal variations can appear as different prosodic scales, which contain the range of within-category variations. The current experiment was intended to show how the prosodic scale corresponding to the range of within-category differences relates to f0 contours for speakers of two Korean dialects, North Kyungsang and South Cholla. In an identification task, participants responded by selecting an item from two answer choices. The probability of choosing the correct response from the two choices was computed by a logistic regression analysis using intercepts and slopes. That is, the correct response between two choices was used to show a linear line with an s-shape presentation. In this paper, to investigate the graded internal structure of labeling, 25%, 50%, and 75% of predicted probability were assessed. Listeners from North Kyungsang showed progressive variations, whereas listeners from South Cholla revealed random patterns in the internal structure of the identification function. In this paper, the results were plotted using scatterplot graphs, applying the range of within-category variation and predicted probability obtained from the logistic regression analyses. The scatterplot graphs showed the different degree of the responses for f0 scales (i.e., variations within categories). The results demonstrate that the gradient structures of native pitch accent users become more progressive in response to f0 scales.

Durational aspects of Korean nasal geminates

  • Oh, Eunhae
    • Phonetics and Speech Sciences
    • /
    • v.9 no.4
    • /
    • pp.19-25
    • /
    • 2017
  • The current study focused on the production of geminate nasal consonants across different word boundary types in Korean as a function of speech style to investigate whether temporal properties are preserved across varying speaking rates. Assimilated geminates in Korean, known as true geminates, are produced with distinctively longer consonant duration compared to singletons. Despite a large body of literature for geminates across different languages, geminates in Korean have been relatively less investigated with respect to the durational patterns in relative terms and temporal variabilities. In this study, singletons, word-internal geminates and word-boundary (fake) geminates produced by ten native Seoul Korean speakers were compared in terms of absolute consonant closure duration, preceding vowel duration, the relative ratios (consonant-to-preceding vowel duration) as well as the temporal variabilities in speech production. The results showed that word-internal geminates were produced with longer consonant duration and greater temporal variabilities than singletons and word-boundary geminates in absolute duration, indicating relatively greater flexibility in timing. However, only word-internal geminates were produced with distinctively longer consonant duration with significantly lower variability in relative duration regardless of speech styles. The results provide some insight into the representation of temporal information in the production of Korean geminate consonants.

Recognition Time Reduction Technique for the Time-synchronous Viterbi Beam Search (시간 동기 비터비 빔 탐색을 위한 인식 시간 감축법)

  • 이강성
    • The Journal of the Acoustical Society of Korea
    • /
    • v.20 no.6
    • /
    • pp.46-50
    • /
    • 2001
  • This paper proposes a new recognition time reduction algorithm Score-Cache technique, which is applicable to the HMM-base speech recognition system. Score-Cache is a very unique technique that has no other performance degradation and still reduces a lot of search time. Other search reduction techniques have trade-offs with the recognition rate. This technique can be applied to the continuous speech recognition system as well as the isolated word speech recognition system. W9 can get high degree of recognition time reduction by only replacing the score calculating function, not changing my architecture of the system. This technique also can be used with other recognition time reduction algorithms which give more time reduction. We could get 54% of time reduction at best.

  • PDF