Search | Korea Science

Speech Recognition Using Noise Robust Features and Spectral Subtraction (잡음에 강한 특징 벡터 및 스펙트럼 차감법을 이용한 음성 인식)

Shin, Won-Ho;Yang, Tae-Young;Kim, Weon-Goo;Youn, Dae-Hee;Seo, Young-Joo
- The Journal of the Acoustical Society of Korea
- /
- v.15 no.5
- /
- pp.38-43
- /
- 1996
This paper compares the recognition performances of feature vectors known to be robust to the environmental noise. And, the speech subtraction technique is combined with the noise robust feature to get more performance enhancement. The experiments using SMC(Short time Modified Coherence) analysis, root cepstral analysis, LDA(Linear Discriminant Analysis), PLP(Perceptual Linear Prediction), RASTA(RelAtive SpecTrAl) processing are carried out. An isolated word recognition system is composed using semi-continuous HMM. Noisy environment experiments usign two types of noises:exhibition hall, computer room are carried out at 0, 10, 20dB SNRs. The experimental result shows that SMC and root based mel cepstrum(root_mel cepstrum) show 9.86% and 12.68% recognition enhancement at 10dB in compare to the LPCC(Linear Prediction Cepstral Coefficient). And when combined with spectral subtraction, mel cepstrum and root_mel cepstrum show 16.7% and 8.4% enhanced recognition rate of 94.91% and 94.28% at 10dB.
PDF

Comparison of Vowel and Text-Based Cepstral Analysis in Dysphonia Evaluation (발성장애 평가 시 /a/ 모음연장발성 및 문장검사의 켑스트럼 분석 비교)

Kim, Tae Hwan;Choi, Jeong Im;Lee, Sang Hyuk;Jin, Sung Min
- Journal of the Korean Society of Laryngology, Phoniatrics and Logopedics
- /
- v.26 no.2
- /
- pp.117-121
- /
- 2015
Background : Cepstral analysis which is obtained from Fourier transformation of spectrum has been known to be effective indicator to analyze the voice disorder. To evaluate the voice disorder, phonation of sustained vowel /a/ sound or continuous speech have been used but the former was limited to capture hoarseness properly. This study is aimed to compare the effectiveness in analysis of cepstrum between the sustained vowel /a/ sound and continuous speech. Methods : From March 2012 to December 2014, total 72 patients was enrolled in this study, including 24 unilateral vocal cord palsy, vocal nodule and vocal polyp patients, respectively. The entire patient evaluated their voice quality by VHI (Voice Handicap Index) before and after treatment. Phonation of sustained vowel /a/ sample and continuous speech using the first sentence of autumn paragraph was subjected by cepstral analysis and compare the pre-treatment group and post-treatment group. Results : The measured values of pre and post treatment in CPP-a (cepstral peak prominence in /a/ vowel sound) was 13.80, 13.91 in vocal cord palsy, 16.62, 17.99 in vocal cord nodule, 14.19, 18.50 in vocal cord polyp respectively. Values of CPP-s (cepstral peak prominence in text-based speech) in pre and post treatment was 11.11, 12.09 in vocal cord palsy, 12.11, 14.09 in vocal cord nodule, 12.63, 14.17 in vocal cord polyp. All 72 patients showed subjective improvement in VHI after treatment. CPP-a showed statistical improvement only in vocal polyp group, but CPP-s showed statistical improvement in all three groups (p<0.05). Conclusion : In analysis of cepstrum, text-based analysis is more representative in voice disorder than vowel sound speech. So when the acoustic analysis of voice by cepstrum, both phonation of sustained vowel /a/ sound and text based speech should be performed to obtain more accurate result.
PDF

Diagnosis of Bearing System using Minimum Variance Cepstrum

Lee, Jeong-Han;Choi, Young-Chul;Park, Jin-Ho;Lee, Won-Hyung;Kim, Chan-Joong
- Proceedings of the Korean Nuclear Society Conference
- /
- 2005.05a
- /
- pp.1255-1256
- /
- 2005
PDF

Cepstrum Analysis of Terrestrial Impact Crater Records

Chang, Heon-Young;Han, Cheong-Ho
- Journal of Astronomy and Space Sciences
- /
- v.25 no.2
- /
- pp.71-76
- /
- 2008
Study of terrestrial impact craters is important not only in the field of the solar system formation and evolution but also of the Galactic astronomy. The terrestrial impact cratering record recently has been examined, providing short- and intermediate-term periodicities, such as, ${\sim}26$ Myrs, ${\sim}37$ Myrs. The existence of such a periodicity has an implication in the Galactic dynamics, since the terrestrial impact cratering is usually interpreted as a result of the environmental variation during solar orbiting in the Galactic plane. The aim of this paper is to search for a long-term periodicity with a novel method since no attempt has been made so far in searching a long-term periodicity in this research field in spite of its great importance. We apply the cepstrum analysis method to the terrestrial impact cratering record for the first time. As a result of the analysis we have found noticeable peaks in the Fourier power spectrum appear ing at periods of ${\sim}300$ Myrs and ${\sim}100$ Myrs, which seem in a simple resonance with the revolution period of the Sun around the Galactic center. Finally we briefly discuss its implications and suggest theoretical study be pursued to explain such a long-term periodicity.
https://doi.org/10.5140/JASS.2008.25.2.071 인용 PDF KSCI

Comparison of MEL-LPC and LPC-MEL Analysis Method for the Korean Speech Recognition Systems. (한국어 음성 인식 시스템을 위한 MEL-LPC 분석 방법과 LPC-MEL 분석 방법의 비교)

김주곤;김범국;정호열;정현열
- Proceedings of the IEEK Conference
- /
- 2001.09a
- /
- pp.833-836
- /
- 2001
본 논문에서는 한국어 음성인식 시스템의 성능 향상을 위해 청각 주파수 분해능을 가진 MEL-LPC Cepstrum을 음소단위의 HMM(Hidden Markov Model)을 기반으로 하는 인식 시스템에 적용하여 그 결과를 비교 검토하였다. 선형예측(LP) 분석 후에 후처리로서 주파수를 왜곡시킨 LPC-MEL 분석이 계산량이 적고 효과적이라 일반적으로 많이 사용되고 있으나 주파수 분해능은 많이 개선되지 않는다. 따라서 본 논문에서는 주파수 분해능을 개선하기 위해, 원 음성신호로부터 직접적으로 멜주파수로 왜곡시킨 후 선형 예측 분석을 수행하는 MEL-LPC 분석방법을 이용한 음소기반의 화자 독립 음성인식 시스템을 구성하여 기존의 LPC-MEL 분석방법과 비교실험을 통하여 MEL-LPC 분석방법의 유효성을 검토하였다. 실험에 사용한 음성 데이터베이스는 음소 및 단어 인식실험에서는 ETRI 445단어 DB, 연속 숫자음인식 실험에서는 KLE 4연속 숫자음 DB를 사용하였다. 화자 독립 음소인식 실험의 경우, 묵음을 제외한 47개의 유사 음소에 대하여 4상태 3출력의 Left-to-Right 모델을이용하였다. 단어 및 연속 숫자음 인식 실험의 경우, 유한상태 네트워크에 의한 OPDP법을 이용하였다. 화자 독립 음소, 단어 및 4연속 숫자음 인식 실험결과, 기존의 LPC-MEL Cepstrum을 사용한 경우보다 MEL-LPC Cepstum을 사용한 경우가 더 높은 인식률을 나타내어 한국어 음성인식 시스템에서 MEL-LPC 분석방법의 유효성을 확인할 수 있었다.
PDF

Temperature Classification of Heat-treated Metals using Pattern Recognition of Ultrasonic Signal (초음파 신호의 패턴 인식에 의한 금속의 열처리 온도 분류)

Im, Rae-Muk;Sin, Dong-Hwan;Kim, Deok-Yeong;Kim, Seong-Hwan
- The Transactions of the Korean Institute of Electrical Engineers A
- /
- v.48 no.12
- /
- pp.1544-1553
- /
- 1999
Recently, ultrasonic testing techniques have been widely used in the evaluation of the quality of metal. In this experiment, six heat-treated temperature of specimen have been considered : 0, 1200, 1250, 1300, 1350 and 1387$^{\circ}C$. As heat-treated temperature increases, the grain size of stainless steel also increases and then, eventually make it destroy. In this paper, a pattern recognition method is proposed to identify the heat-treated temperature of metals by evidence accumulation based on artificial intelligence with multiple feature parameters; difference absolute mean value(DAMV), variance(VAR), mean frequency(MEANF), auto regressive model coefficient(ARC), linear cepstrum coefficient(LCC) and adaptive cepstrum vector(ACV). The grain signal pattern recognition is carried out through the evidence accumulation procedure using the distances measured with reference parameters. Especially ACV is superior to the other parameters. The results (96% successful pattern classification) are presented to support the feasibility of the suggested approach for ultrasonic grain signal pattern recognition.
PDF

wheelchair system design on speech recognition function (음성인식 기능을 탑재한 다기능 휠체어 시스템 설계 및 구현)

김정훈;류홍석;강재명;강성인;김관형;이상배
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 2002.05a
- /
- pp.1-5
- /
- 2002
The purpose of this paper is developing a speech recognition module in a wheelchair for the sake of convenience. of the disability. For this system, we used TMS320C32 as a main processor; eliminated noise by applying Winer filler while considering characteristics of noise environment in pre-processing stage, and; extracted 12 feature patterns per france using LPC&Cepstrum. Then, we implemented the hybrid form combining DTW (Dynamic Time Warping), which is generally used for isolated words in the conventional algorithms, in the recognition Part, and NN (Neural network) to prevent any error of recognition. In this research, we achieved a recognition rate of more than 96% on isolated words when DTW and Hybrid forms were individually experimented in noise environment
PDF

Automatic Detection of Cow's Oestrus in Audio Surveillance System

Chung, Y.;Lee, J.;Oh, S.;Park, D.;Chang, H.H.;Kim, S.
- Asian-Australasian Journal of Animal Sciences
- /
- v.26 no.7
- /
- pp.1030-1037
- /
- 2013
Early detection of anomalies is an important issue in the management of group-housed livestock. In particular, failure to detect oestrus in a timely and accurate way can become a limiting factor in achieving efficient reproductive performance. Although a rich variety of methods has been introduced for the detection of oestrus, a more accurate and practical method is still required. In this paper, we propose an efficient data mining solution for the detection of oestrus, using the sound data of Korean native cows (Bos taurus coreanea). In this method, we extracted the mel frequency cepstrum coefficients from sound data with a feature dimension reduction, and use the support vector data description as an early anomaly detector. Our experimental results show that this method can be used to detect oestrus both economically (even a cheap microphone) and accurately (over 94% accuracy), either as a standalone solution or to complement known methods.
https://doi.org/10.5713/ajas.2012.12628 인용 PDF KSCI

An Enhanced Text-Prompt Speaker Recognition Using DTW (DTW를 이용한 향상된 문맥 제시형 화자인식)

신유식;서광석;김종교
- The Journal of the Acoustical Society of Korea
- /
- v.18 no.1
- /
- pp.86-91
- /
- 1999
This paper presents the text-prompt method to overcome the weakness of text-dependent and text-independent speaker recognition. Enhanced dynamic time warping for speaker recognition algorithm is applied. For the real-time processing, we use a simple algorithm for end-point detection without increasing computational complexity. The test shows that the weighted-cepstrum is most proper for speaker recognition among various speech parameters. As the experimental results of the proposed algorithm for three prompt words, the speaker identification error rate is 0.02%, and when the threshold is set properly, false rejection rate is 1.89%, false acceptance rate is 0.77% and verification total error rate is 0.97% for speaker verification.
PDF

HMM-based Speech Recognition using DMS Model and Double Spectral Feature (DMS 모델과 이중 스펙트럼 특징을 이용한 HMM에 의한 음성 인식)

Ann Tae-Ock
- Journal of the Korea Academia-Industrial cooperation Society
- /
- v.7 no.4
- /
- pp.649-655
- /
- 2006
This paper proposes a HMM-based recognition method using DMSVQ(Dynamic Multi-Section Vector Quantization) codebook by DMS model and double spectral feature, as a method on the speech recognition of speaker-independent. LPC cepstrum parameter is used as a instantaneous spectral feature and LPC cepstrum's regression coefficient is used as a dynamic spectral feature These two spectral features are quantized as each VQ codebook. HMM using DMS model is modeled by receiving instantaneous spectral feature and dynamic spectral feature by input. Other experiments to compare with the results of recognition experiments using proposed method are implemented by the various conventional recognition methods under the equivalent environment of data and conditions. Through the experiment results, it is proved that the proposed method in this paper is superior to the conventional recognition methods.
PDF

Search Result 274, Processing Time 0.029 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)