• Title/Summary/Keyword: 음성검출기

Search Result 137, Processing Time 0.022 seconds

A Study on Glottal Spectrum Analysis According to the Distance between the Microphone and the lips (Microphone 거리에 따른 Glottal Spectrum 성분 분석에 관한 연구)

  • Park Hyunyoung;Jang Kyunga;Bae Myungjin
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • spring
    • /
    • pp.65-68
    • /
    • 2002
  • 현재 음성인식기는 다 채널의 음성입력방식을 사용하고 있는 추세이다. 이런 방법으로 음성인식기를 사용할 때에 자동적으로 음성을 검출하는 음성입력 방식은 발성자와 마이크간의 거리에 따라 Glottal Spectrum 성분이 변하는 특성을 가지고 있다. 이러한 Glottal Spectrum 성분은 a=R1/R0 (LPC 포락선의 기울기) 로 나타낼 수 있다. 본 논문에서는 발성자와 마이크 거리에 따른 Glottal Spectrum 성분을 비교 분석 하고자 한다.

  • PDF

Frequency Domain Double-Talk Detector Based on Gaussian Mixture Model (주파수 영역에서의 Gaussian Mixture Model 기반의 동시통화 검출 연구)

  • Lee, Kyu-Ho;Chang, Joon-Hyuk
    • The Journal of the Acoustical Society of Korea
    • /
    • v.28 no.4
    • /
    • pp.401-407
    • /
    • 2009
  • In this paper, we propose a novel method for the cross-correlation based double-talk detection (DTD), which employing the Gaussian Mixture Model (GMM) in the frequency domain. The proposed algorithm transforms the cross correlation coefficient used in the time domain into 16 channels in the frequency domain using the discrete fourier transform (DFT). The channels are then selected into seven feature vectors for GMM and we identify three different regions such as far-end, double-talk and near-end speech using the likelihood comparison based on those feature vectors. The presented DTD algorithm detects efficiently the double-talk regions without Voice Activity Detector which has been used in conventional cross correlation based double-talk detection. The performance of the proposed algorithm is evaluated under various conditions and yields better results compared with the conventional schemes. especially, show the robustness against detection errors resulting from the background noises or echo path change which one of the key issues in practical DTD.

Optimization of Detection Method Using a Moving Average Estimator for Speech Enhancement (음성강화를 위한 이동 평균 예측량 기반의 검출방법 최적화)

  • Lee, Soo-Jeong;Shin, Kye-Hyeon;Kim, Soon-Hyob
    • Journal of the Institute of Electronics Engineers of Korea SP
    • /
    • v.44 no.3
    • /
    • pp.97-104
    • /
    • 2007
  • Adaptive echo canceller(AEC) has become an important component in speech communication systems, including mobile phones and speech recognition. In these applications, the acoustic echo path has a long impulse response. We propose a moving-averge least mean square(MVLMS) algorithm with a detection method for acoustic echo cancellation. Using, the result of the tests that used colored input models clearly shows that the MVLMS detection algorithm has convergence performance superior to the least mean square(LMS) detection algorithm alone. Although the computational complexity of the new MVLMS algorithm is only slightly greater than that of the standard LMS detection algorithm, the new algorithm confers a significant improvement in stability.

Scoring Methods for Improvement of Speech Recognizer Detecting Mispronunciation of Foreign Language (외국어 발화오류 검출 음성인식기의 성능 개선을 위한 스코어링 기법)

  • Kang Hyo-Won;Kwon Chul-Hong
    • MALSORI
    • /
    • no.49
    • /
    • pp.95-105
    • /
    • 2004
  • An automatic pronunciation correction system provides learners with correction guidelines for each mispronunciation. For this purpose we develope a speech recognizer which automatically classifies pronunciation errors when Koreans speak a foreign language. In order to develope the methods for automatic assessment of pronunciation quality, we propose a language model based score as a machine score in the speech recognizer. Experimental results show that the language model based score had higher correlation with human scores than that obtained using the conventional log-likelihood based score.

  • PDF

Pronunciation Network Construction of Speech Recognizer for Mispronunciation Detection of Foreign Language (한국인의 외국어 발화오류 검출을 위한 음성인식기의 발음 네트워크 구성)

  • Lee Sang-Pil;Kwon Chul-Hong
    • MALSORI
    • /
    • no.49
    • /
    • pp.123-134
    • /
    • 2004
  • An automatic pronunciation correction system provides learners with correction guidelines for each mispronunciation. In this paper we propose an HMM based speech recognizer which automatically classifies pronunciation errors when Koreans speak Japanese. We also propose two pronunciation networks for automatic detection of mispronunciation. In this paper, we evaluated performances of the networks by computing the correlation between the human ratings and the machine scores obtained from the speech recognizer.

  • PDF

Automatic Detection of Mispronunciation Using Phoneme Recognition For Foreign Language Instruction (음성인식기를 이용한 한국인의 외국어 발화오류 자동 검출)

  • Kwon Chul-Hong;Kang Hyo-Won;Lee Sang-Pil
    • MALSORI
    • /
    • no.48
    • /
    • pp.127-139
    • /
    • 2003
  • An automatic pronunciation correction system provides learners with correction guidelines for each mispronunciation. In this paper we propose an HMM based speech recognizer which automatically classifies pronunciation errors when Korean speak Japanese. For this purpose we also develop phoneme recognizers for Korean and Japanese. Experimental results show that the machine scores of the proposed recognizer correlate with expert ratings well.

  • PDF

Speech Activity Decision with Lip Movement Image Signals (입술움직임 영상신호를 고려한 음성존재 검출)

  • Park, Jun;Lee, Young-Jik;Kim, Eung-Kyeu;Lee, Soo-Jong
    • The Journal of the Acoustical Society of Korea
    • /
    • v.26 no.1
    • /
    • pp.25-31
    • /
    • 2007
  • This paper describes an attempt to prevent the external acoustic noise from being misrecognized as the speech recognition target. For this, in the speech activity detection process for the speech recognition, it confirmed besides the acoustic energy to the lip movement image signal of a speaker. First of all, the successive images are obtained through the image camera for PC. The lip movement whether or not is discriminated. And the lip movement image signal data is stored in the shared memory and shares with the recognition process. In the meantime, in the speech activity detection Process which is the preprocess phase of the speech recognition. by conforming data stored in the shared memory the acoustic energy whether or not by the speech of a speaker is verified. The speech recognition processor and the image processor were connected and was experimented successfully. Then, it confirmed to be normal progression to the output of the speech recognition result if faced the image camera and spoke. On the other hand. it confirmed not to output of the speech recognition result if did not face the image camera and spoke. That is, if the lip movement image is not identified although the acoustic energy is inputted. it regards as the acoustic noise.

A study on real-time implementation of speech recognition and speech control system using dSPACE board (dSPACE 보드를 이용한 음성인식 명령처리시스템 실시간 구현에 관한 연구)

  • 김재웅;정원용
    • Proceedings of the Korea Institute of Convergence Signal Processing
    • /
    • 2000.12a
    • /
    • pp.173-176
    • /
    • 2000
  • 음성은 인간이 가진 가장 편리한 제어전송수단으로 이를 통한 제어는 인간에게 많은 편리함을 제공할 것이다. 본 논문에서는 다층구조 신경망(Multi-Layer Perceptron)을 이용하여 간단한 음성인식 명령처리시스템을 Matlab 상에서 구성해 보았다. 음성인식을 통한 제어의 목적을 위해 화자종속, 고립단어인식기를 목표로 설정하여 연구를 수행하였다. 음성의 시작점과 끝점을 검출하기 위해 단구간 에너지와 영교차율(ZCR)을 이용하였고 인식기의 특징파라미터로는 12차 LPC켑스트럼 계수를 사용하였다. 그리고 신경망의 출력값을 기동, 정지시에 활성화되도록 3개의 계층으로 하였고, 신경망의 뉴런의 개수를 각각 12, 12, 2으로 설정하였다. 먼저 기준음성패턴으로 학습시킨 후에 Matlab 환경하에 동작하는 dSPACE 실시간처리보드에 변환된 C프로그램을 다운로드하고, 음성을 입력하여 인식 후 dSPACE보드의 D/A컨버터의 출력단에 연결된 DC모터를 기동, 정지제어를 수행하였다. 실시간 음성인식 명령처리 시스템 구현을 통하여 원격제어와 같은 음성명령을 통한 제어가 가능함을 확인할 수 있었다.

  • PDF

Speaker Adaptation Performance Evaluation in Keyword Spotting System (500단어급 핵심어 검출기에서 화자적응 성능 평가)

  • Seo Hyun-Chul;Lee Kyong-Rok;Kim Jin-Young;Choi Seung-Ho
    • MALSORI
    • /
    • no.43
    • /
    • pp.151-161
    • /
    • 2002
  • This study presents performance analysis results of speaker adaptation for keyword spotting system. In this paper, we implemented MLLR (Maximum Likelihood Linear Regression) method on our middle size vocabulary keyword spotting system. This system was developed for directory services of universities and colleges. The experimental results show that speaker adaptation reduces the false alarm rate to 1/3 with the preservation of the mis-detection ratio. This improvement is achieved when speaker adaptation is applied to not only keyword models but also non-keyword models.

  • PDF

기능적 자기공명영상 및 확산텐서영상을 이용한 전음성 난청과 감각신경성 난청군의 비교 연구: 예비 결과

  • 이재준;황문정;이영주;김인성;배성진;장용민;이상흔;우성구;강덕식
    • Proceedings of the KSMRM Conference
    • /
    • 2003.10a
    • /
    • pp.94-94
    • /
    • 2003
  • 목적: 기능적 자기공명영상과 확산텐서영상기법을 이용하여 전음성 난청과 감각신경성 난청에서의 뇌활성화 양상 그리고 청신경경로상의 차이점을 비교 연구하고자 하였다. 대상 및 방법: 전음성 난청군 (n=4)과 감각신경성 난청군(n=5) 그리고 정상군(n=5)에서의 기능적 자기 공명영상과 확산텐서영상을 획득하였다. 기능적 자기공명영상의 경우 1.5T Siemens MR scanner에서 BOLD 기법을 이용하여 500 Hz 순음 청각자극에 대한 뇌활성화 영역을 검출하였고 영상촬영시 발생하는 기계적 소음을 차폐하기 위한 청각자극기를 특별히 제작하여 사용하였다. 뇌백질신경로를 영상화하는 확산텐서영상은 3.0T GE whole body MR scanner를 사용하였으며 미세한 확산운동을 검출하기 위해 초고속 영상기법인 EPI 기법을 사용하였다. 영상의 화질을 높이기 위해 공간적으로 25개의 다른 방향으로 확산경사자장을 가하였다. 청신경로의 비등방성 영상, 신경로 방향 영상등을 구현하기 위해 획득한 확산영상들에 대한 영상 후처리과정을 시행하였다.

  • PDF